Data Center Operations Project Manager
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the Role
We are seeking a Data Center Real Estate & Development Specialist to own our data center capacity end to end, from sourcing space and power to delivering operational facilities. You will source candidate sites through brokers and landlords nationwide, running technical and commercial diligence, negotiating leases, and readying each facility for GPU deployment. You will be the central point of coordination across brokers, landlords, utilities, engineering firms, contractors, OEM partners, and internal stakeholders, keeping acquisition, design, permitting, procurement, and construction on schedule and on budget.
This role is ideal for someone who thrives operating with urgency, is comfortable getting on a plane to walk a site anywhere in the United States, can make sound judgment calls under time pressure, and can hold multiple vendors accountable to a shared, aggressive timeline.
Key Responsibilities
Site Acquisition & Real Estate
Build and manage a nationwide network of commercial real estate brokers, landlords, and owners to source space-and-power opportunities (powered shell, industrial conversions, existing data center capacity), maintaining a pipeline of on-market and off-market candidate sites.
Travel across the United States to tour candidate facilities and lead on-site diligence, walkthroughs, and inspections through deal close.
Lead commercial negotiations end to end: LOIs, term sheets, and lease agreements with landlords, partnering with legal counsel and finance on rent structure, tenant improvements, power terms, expansion rights, and local incentives.
Build comparative site economics (all-in $/kW, upgrade CapEx, lease terms) and present clear go/no-go recommendations to leadership.
Technical Diligence & Deployment Fit
Run technical diligence on candidate sites: utility capacity and interconnection status, electrical infrastructure (transformers, medium-voltage step-down, switchgear, busways, PDUs, UPS, generators), cooling and water availability, structural and floor loading, fiber connectivity, zoning and permitting posture, and expansion potential.
Coordinate with OEM and hardware partners to validate that target facilities meet future deployment requirements (rack density, power per rack, liquid cooling readiness, floor loading, delivery logistics) before we commit.
Facility Conversion & Build Execution
Coordinate directly with our electrical infrastructure partner on power distribution and cooling build-out scope, including UPS, PDUs, switchgear, ATS, CDUs, chilled water systems, and in-row cooling, ensuring designs align with project timeline and budget.
Scope and manage facility improvement projects with general contractors, engineering firms, utilities, municipalities, and equipment vendors, covering upgrades such as new transformers, switchgear, busways, and cooling infrastructure.
Own permitting and AHJ (Authority Having Jurisdiction) coordination, securing approvals without delaying the critical path.
Drive design and engineering review cycles, tracking open technical questions, decisions, and dependencies across electrical, mechanical, and structural workstreams.
Coordinate procurement of long-lead equipment, tracking order timing, lead times, and budget against project CapEx.
Manage relationships with server/GPU hardware vendors on rack infrastructure, delivery schedules, and integration with facility infrastructure and deployment plans.
Manage the master project schedule across parallel tracks and sites, identifying critical path items and escalating risks early.
Manage site logistics, including site visits, inspections, and access for partners and contractors.
Stakeholder Management & Reporting
Serve as the primary point of contact for all external partners (brokers, landlords, electrical infrastructure partner, engineering firms, contractors, utilities, hardware vendors) and internal stakeholders (leadership, finance).
Track and report pipeline and project status, risks, and decisions needed to leadership on a regular cadence.
Must-Haves
5+ years in data center development, site selection, or large-scale infrastructure/construction project management, ideally spanning both deal-making and build execution.
Demonstrated experience sourcing industrial or data center properties through broker networks and direct landlord relationships.
Experience negotiating commercial real estate transactions (LOIs, leases), ideally for industrial or mission-critical facilities, working alongside legal counsel.
Willingness to travel extensively across the United States (up to 50%, often on short notice) for tours, diligence, and site oversight.
Working knowledge of data center power and cooling systems (utility service, transformers and step-down distribution, switchgear, busways, PDUs, UPS, CDUs, liquid cooling) sufficient to have informed technical conversations with vendors and engineers.
Experience with permitting and AHJ processes for industrial or commercial construction projects.
Experience coordinating electrical contractors, engineering firms, general contractors, and major equipment vendors.
Strong project scheduling skills (critical path method, MS Project, Smartsheet, or similar tools).
Strong analytical skills for site comparison, deal modeling, and CapEx budgeting.
Ability to manage multiple concurrent workstreams and vendors under aggressive timelines; excellent written and verbal communication across technical and non-technical stakeholders.
Nice-to-Haves
Prior experience with AI/HPC data center builds or GPU cluster deployments.
Direct utility coordination experience: capacity studies, interconnection agreements, or tariff discussions.
Familiarity with liquid cooling infrastructure (direct-to-chip, CDUs) and high-density power distribution (480V, 415V step-down systems).
Experience with facility conversions (brownfield) rather than greenfield builds.
Background at a colocation provider, hyperscaler, or data center developer.
Bachelor's degree in engineering, construction management, real estate, or a related field; PMP certification.
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Model Implementation Engineer
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
We are seeking a highly skilled Model Implementation Engineer who is passionate about bringing cutting-edge machine learning models into production-ready systems. In this role, you will implement, maintain, and optimize a large and evolving library of state-of-the-art models across modalities, ensuring high performance and reliability from day one.
You will work at the intersection of research and systems, translating the latest ideas into robust, scalable implementations. This includes collaborating closely with GPU kernel and systems teams to ensure models are efficiently executed on modern accelerators.
This role is ideal for someone who thrives in fast-moving environments, enjoys working across a wide range of model architectures, and wants to play a key role in enabling rapid adoption of the latest advancements in AI.
Key Responsibilities
Maintain and evolve a large-scale library of modern machine learning models, including but not limited to LLMs, ASR, TTS, image and video models, and diffusion-based systems.
Implement new model architectures and research ideas, ensuring correctness, scalability, and production readiness.
Rapidly integrate newly released open-source models to enable day-0 support across the platform.
Collaborate closely with GPU kernel and systems teams to optimize model execution and improve overall performance.
Benchmark models rigorously and ensure they meet internal performance, latency, and efficiency standards.
Contribute to the canonicalization and standardization of model implementations across the library.
Develop and maintain internal tooling, testing frameworks, and documentation to support model reliability and reproducibility.
Must-Haves
At least 3 years of industry or research experience in model implementation or applied machine learning.
Master of Science (or higher) in Computer Science, Machine Learning, Electrical Engineering, Applied Mathematics, or a related field.
Strong programming skills in Python and experience working with modern ML frameworks.
Hands-on experience with JAX and/or PyTorch (JAX strongly preferred).
Proven experience maintaining and developing model libraries or reusable ML components.
Solid understanding of deep learning architectures across multiple domains (e.g., NLP, vision, speech, generative models).
Experience implementing models from research papers and adapting them for real-world usage.
Ability to work across teams and collaborate with systems and performance engineering groups
Nice-to-Haves
Experience with model performance optimization and profiling.
Familiarity with low-level performance considerations when running models on GPUs/TPUs.
Experience working with large-scale model training or inference systems.
Contributions to open-source model repositories or ML frameworks.
Experience with JAX-first workflows and advanced features (e.g., pjit, xmap, or custom transformations).
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
LLM Dataset Engineer
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
Role Overview
Sciforium is seeking a highly technical and visionary LLM Dataset Engineer to lead the strategy, creation, and curation of the massive datasets that power our foundation models. We believe that in the era of LLMs, data is the primary competitive advantage. In this role, you will own the end-to-end data lifecycle—from raw web-scale crawling to the fine-grained human-alignment datasets that define model behavior.
This position is ideal for a scientist who views data as a high-scale engineering challenge and an analytical puzzle. You will not just "provide" data; you will design the taxonomies, filtering heuristics, and post-training pipelines that ensure our models are world-class in reasoning, safety, and multimodal understanding.
Key Responsibilities
Foundation Dataset Strategy: Own the end-to-end creation of pre-training datasets for LLMs. This includes defining the mix of web data, code, books, and technical papers to optimize for downstream model performance.
Petabyte-Scale Curation: Design and implement sophisticated pipelines for data cleaning, exact/fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
Post-Training & Alignment Data: Lead the development of high-quality post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data (RLHF/DPO).
Multimodal Expansion: Drive the acquisition and processing of vision and video data, navigating the complexities of multimodal alignment, video compression, and temporal data consistency.
High-Performance Engineering: Develop high-throughput data processing scripts using Python, leveraging multiprocessing and multithreading to handle massive-scale ingestion and transformation without bottlenecks.
Data Profiling & Analysis: Conduct deep-dive statistical analysis on training corpora to identify biases, gaps in knowledge, and quality regressions, ensuring the "diet" of the model is mathematically balanced.
Synthetic Data Generation: (Added Value) Design pipelines to generate high-reasoning synthetic data to augment gaps in natural datasets, utilizing existing models for data labeling and refinement.
Must-Haves
5+ years of industry experience in Data Science or Machine Learning, with a proven track record of building and managing datasets for foundation models.
Deep Proficiency in Python: Expert-level skills with a focus on high-performance code, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.
Petabyte-Scale Experience: Demonstrated experience working with petabyte-scale datasets that have been directly used to train production-grade LLMs or Large Vision Models.
Dataset Reconstruction: Experience building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data.
Post-Training Expertise: Hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following, including the management of human-labeling workflows and quality gold-sets.
Data Tooling: Mastery of data-at-scale frameworks such as Spark, Ray, or high-performance data-loading formats (e.g., WebDataset, Parquet).
Nice-to-Haves
Computer Vision (CV) Curation: Experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines).
Multimodal Crawling: Familiarity with large-scale crawling of multimodal data and the associated challenges of video processing, codecs, and compression.
Taxonomy Design: Experience in designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
Research Background: A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
LLM Training Engineer
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the Role
As an LLM Training Engineer, you’ll work across the full foundation-model stack: pretraining and scaling, post-training and Reinforcement Learning, sandbox environments for evaluation and agentic learning, and deployment + inference optimization. You’ll build and iterate quickly on research ideas, contribute production-grade infrastructure, and help deliver models that can serve real-world use cases at scale.
What you’ll work on
This role spans multiple tracks - candidates may focus on one or contribute across several. Examples include:
Pretraining & Scaling
Train large byte-native foundation models across massive, heterogeneous corpora
Design stable training recipes and scaling laws for novel architectures
Improve throughput, memory efficiency, and utilization on large GPU clusters
Build and maintain distributed training infrastructure and fault-tolerant pipelines
Post-training & RL
Develop post-training pipelines (SFT, preference optimization, RLHF/RLAIF, RL)
Curate and generate targeted datasets to improve specific model capabilities
Build reward models and evaluation frameworks to drive iterative improvement
Explore inference-time learning and compute techniques to enhance performance
Sandbox Environments & Evaluation
Build scalable sandbox environments for agent evaluation and learning
Create realistic, high-signal automated evals for reasoning, tool use, and safety
Design offline + online environments that support RL-style training at scale
Instrument environments for observability, reproducibility, and iteration speed
Deployment & Inference Optimization
Optimize inference throughput/latency for byte-native architectures
Build high-performance serving pipelines (KV caching, batching, quantization, etc.)
Improve end-to-end model efficiency, cost, and reliability in production
Profile and optimize GPU kernels, runtime bottlenecks, and memory behavior
Ideal candidate credentials
Technical strength
Strong general software engineering skills (writing robust, performant systems)
Experience with training or serving large neural networks (LLMs or similar)
Solid grasp of deep learning fundamentals and modern literature
Comfort working in high-performance environments (GPU, distributed systems, etc.)
Relevant experience (one or more)
Pretraining / large-scale distributed training (FSDP/ZeRO/Megatron-style systems)
Post-training pipelines (SFT, RLHF/RLAIF, preference optimization, eval loops)
Building RL environments, simulators, or agent frameworks
Inference optimization, model compression, quantization, kernel-level profiling
Building large ETL pipelines for internet-scale data ingestion and cleaning
Owning end-to-end production ML systems with monitoring and reliability
Research orientation
Ability to propose and evaluate research ideas quickly
Strong experimental hygiene: ablations, metrics, reproducibility, analysis
Bias toward building — you can turn ideas into working code and results
Education
MS or PhD in Computer Science, Machine Learning, AI, Mathematics, or related field
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Senior Research Scientist
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
Sciforium is seeking an exceptional Senior Research Scientist specializing in advanced AI/ML to push the boundaries of scientific and engineering innovation. In this role, you will lead cutting-edge research initiatives in large language models, generative media, model architecture, optimization, and scalable training systems. You will work hands-on with modern ML frameworks, contribute original research, and collaborate with engineering teams to bring impactful models into production. This role is ideal for a highly motivated researcher who is passionate about driving foundational breakthroughs in AI.
What you'll do
Lead research in advanced machine learning areas such as LLMs, generative AI, foundational modeling, optimization techniques, diffusion models, and novel Transformer architectures.
Design, implement, and evaluate novel ML algorithms using frameworks like PyTorch and JAX.
Conduct large-scale distributed training experiments using multi-GPU/TPU systems and modern compute infrastructure.
Drive performance improvements through framework debugging, speed optimization, and training pipeline enhancements.
Produce high-quality research output including papers, internal reports, patents, and reproducible code.
Collaborate with engineering and product teams to translate research prototypes into scalable production systems.
Stay ahead of the latest research developments and integrate state-of-the-art techniques into Sciforium’s AI roadmap.
Mentor junior researchers and contribute to building a world-class AI research culture.
Ideal candidate profile
PhD in Computer Science, Machine Learning, AI, Mathematics, or a related field (required).
5+ years of academic or industry research experience (with flexibility for exceptional fresh PhDs from top programs).
Proven track record of impactful research, evidenced by publications in NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, or similar top-tier venues.
Strong coding proficiency in Python, with deep experience in PyTorch and/or JAX.
Experience with debugging, performance profiling, and speed optimization.
Expertise in distributed training at scale.
Nice-to-have
Extensive experience with JAX/Flax/XLA stack
Experience with efficient model serving, inference optimization, model compression, or production ML systems.
Hands-on research or engineering experience with diffusion models or generative modeling.
Experience contributing to product development or collaborating with cross-functional engineering teams.
Open-source contributions to major ML frameworks or libraries.
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Lead Software Engineer, Model Serving Platform
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
This is a rare chance to help architect and lead the development of Sciforium’s next-generation model serving platform, the high-performance engine that will bring a multimodal, highly efficient foundation model to market. As a senior technical leader, you’ll not only build core components yourself but also guide and mentor other engineers, influencing engineering direction, standards, and execution quality.
You will learn and shape the full AI stack: from GPU kernels and quantized execution paths to distributed serving, scheduling, and the APIs that power real-time AI applications. If you enjoy deep systems work, thrive on ownership, and want to lead engineers in building foundational AI infrastructure, this role puts you at the center of SciForium’s mission and growth.
What you'll do
Lead the technical direction of the model serving platform, owning architecture decisions and guiding engineering execution.
Build core serving components including execution runtimes, batching, scheduling, and distributed inference systems.
Develop high-performance C++ and CUDA/HIP modules, including custom GPU kernels and memory-optimized runtimes.
Collaborate with ML researchers to productionize new multimodal models and ensure low-latency, scalable inference.
Build Python APIs and services that expose model capabilities to downstream applications.
Mentor and support other engineers through code reviews, design discussions, and hands-on technical guidance.
Drive performance profiling, benchmarking, and observability across the inference stack.
Ensure high reliability and maintainability through testing, monitoring, and engineering best practices.
Troubleshoot and resolve complex issues across GPU, runtime, and service layers.
Ideal candidate profile
Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
5+ years of experience designing and building scalable, reliable backend systems or distributed infrastructure.
Strong understanding of LLM inference mechanics (prefill vs decode, batching, KV cache)
Experience with Kubernetes/Ray, Containerization
Strong proficiency in C++, Python.
Strong debugging, profiling, and performance optimization skills at the system level.
Ability to collaborate closely with ML researchers and translate model or runtime requirements into production-grade systems.
Effective communication skills and the ability to lead technical discussions, mentor engineers, and drive engineering quality.
Comfortable working from the office and contributing to a fast-moving, high-ownership team culture.
Nice-to-have
Experience with ML systems engineering, distributed GPU scheduling, open source inference engine like vLLM, Sglang, or TRT-LLM
Experience in building large scale ML/MLOps infrastructure
Proficiency in CUDA or ROCm and experience with GPU profiling tools
Experience at an AI/ML startup, research lab, or Big Tech infrastructure/ML team.
Familiarity with multimodal model architectures, raw-byte models, or efficient inference techniques.
Contributions to open-source ML or HPC infrastructure
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Software Engineer, Fullstack
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
This role offers a unique opportunity to work on the core systems that power Sciforium’s multimodal AI models. You’ll help build the model serving platform working across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable engine for real-time AI applications.
You’ll gain hands-on experience with performance engineering, learn how large AI models are optimized and deployed at scale, and collaborate closely with ML researchers and experienced systems engineers. If you enjoy delighting customers with a great Developer Experience, care deeply about performance, and want exposure to the full AI stack, this role provides both high-impact work and strong growth potential.
What you'll do
Interactive AI Chat interface: Design and build a low-latency, chat-like interface for users to test our LLMs.
Developer Console: Build the mission-critical UI where users generate API Keys, set budget limits, and view real-time usage graphs
Payments & Billing: Integrate Stripe (or similar) to handle complex subscription models.
Documentation Portal: creating a dynamic API reference section
CLI/SDK Integration: Write simple client-side wrappers to help users connect to our API endpoints easily.
Backend APIs: Building secure API gateways and implementing low latency streaming solutions.
Ideal candidate profile
Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience)
3+ years of software engineering experience, with a focus on front end development.
Strong proficiency in Typescript, Python.
Understanding of responsive design, and UX fundamentals
Strong collaboration and communication skills, with the ability to work effectively across engineering and ML teams.
Comfortable working from the office and contributing to a fast-moving, high-ownership team culture.
Nice-to-have
Experience with ML systems engineering, open source inference engine like vLLM, Sglang, or TRT-LLM
Streaming Expertise: You understand the difference between WebSockets and Server-Sent Events (SSE) and know how to handle "jittery" network streams without freezing the UI.
Billing integration: Experience integrating Stripe Elements, managing webhooks for payment success/failure, and handling "SaaS" logic (pro-rating, tiers).
Documentation Mindset: You appreciate good DX (Developer Experience). You know how to render an API spec into a readable web page.
Python/Backend Knowledge: You can read a FastAPI or vLLM backend codebase to understand how data is being sent to you
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Distributed Training and Inference Engineer
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
Sciforium is seeking a highly skilled Distributed Training and Inference Engineer to build, optimize, and maintain the critical software stack that powers our large-scale AI training and serving workloads. In this role, you will work across the entire machine learning infrastructure from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch to ensure our distributed training systems are fast, scalable, stable, and efficient.
This position is ideal for someone who loves deep systems engineering, debugging complex hardware–software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.
What you'll do
Software Stack Maintenance: Maintain, update, and optimize critical ML libraries and frameworks including JAX, PyTorch, CUDA, and ROCm across multiple environments and hardware configurations.
End-to-End Stack Ownership: Build, maintain, and continuously improve the entire ML software stack from ROCm/CUDA drivers to high-level JAX/PyTorch tooling.
Distributed System Optimization: Ensure all model implementations are efficiently sharded, partitioned, and configured for large-scale distributed training and serving.
System Integration: Continuously integrate and validate modules for runtime correctness, memory efficiency, and scalability across multi-node GPU/accelerator clusters.
Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.
Debugging & Reliability: Troubleshoot complex hardware–software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.
Collaborate with research, infrastructure, and kernel engineering teams to improve system throughput, stability, and developer experience.
Ideal candidate profile
5+ years of industry experience in ML systems, distributed training, or related fields.
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical fields.
Strong programming experience in Python, C++, and familiarity with ML tooling and distributed systems.
Deep understanding of profiling tools (e.g., Nsight, ROCm Profiler, XLA profiler, TPU tools).
Deep expertise with partitioning configuration on the modern ML frameworks such as PyTorch and JAX.
Experience with multi-node distributed training systems and orchestration frameworks (DTensor, GSPMD, etc.).
Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.
Nice-to-have
Extensive experience with the XLA/JAX stack, including compilation internals and custom lowering paths.
Familiarity with distributed serving or large-scale inference frameworks (e.g., vLLM, TensorRT, FasterTransformer).
Background in GPU kernel optimization or accelerator-aware model partitioning.
Strong understanding of low-level C++ building blocks used in ML frameworks (e.g., XLA, CUDA kernels, custom ops).
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Senior AI Serving Engineer, Backend
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
This role offers a unique opportunity to work on the core systems that power Sciforium’s multimodal AI models. You’ll help build the model serving platform working across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable engine for real-time AI applications.
You’ll gain hands-on experience with performance engineering, learn how large AI models are optimized and deployed at scale, and collaborate closely with ML researchers and experienced systems engineers. If you enjoy low-level programming, care deeply about performance, and want exposure to the full AI stack, this role provides both high-impact work and strong growth potential.
What you'll do
Build the model serving platform, including API, Control Plane, Billing, Monitoring, and distributed inference features.
Collaborate with ML researchers to integrate new multimodal models into production workflows.
Write reliable, maintainable code with strong testing and documentation practices.
Provide operational support for keeping our production services highly performant, available and reliable
Help troubleshoot complex issues across runtime, service, and GPU layers, working closely with other engineers.
Ideal candidate profile
Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience)
3+ years of software engineering experience, with a focus on infrastructure or machine learning systems.
Strong proficiency in C++/Python/Go/Rust
Experience with Kubernetes, Containerization
Experience in building large scale ML/MLOps infrastructure
Strong collaboration and communication skills, with the ability to work effectively across engineering and ML teams.
Comfortable working from the office and contributing to a fast-moving, high-ownership team culture.
Nice-to-have
Experience with ML systems engineering, open source inference engine like vLLM, Sglang, or TRT-LLM
Proficiency in CUDA or ROCm and experience with GPU profiling tools
Contributions to open-source ML or HPC infrastructure
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
GPU Kernel Engineer
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
About the role
We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you will design and optimize custom GPU kernels that power next-generation large-scale AI systems. You will work across the hardware–software stack, from low-level kernel development to integrating optimized ops into high-level ML frameworks used for large-scale training and inference.
This role is ideal for someone who thrives at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads, and who wants to make meaningful contributions to the efficiency and scalability of our ML platform.
Key Responsibilities
Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.
Profile and optimize end-to-end performance of ML operations, with a focus on large-scale LLM training and inference.
Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes.
Develop performance models, identify bottlenecks, and deliver kernel-level improvements that significantly accelerate AI workloads.
Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack.
Work closely with hardware vendors (NVIDIA/AMD) and stay current on the latest GPU architecture capabilities and compiler/toolchain improvements.
Contribute to tooling, documentation, benchmarking suites, and testing frameworks to ensure correctness and performance reproducibility.
Must-Haves
5+ years of industry or research experience in GPU kernel development or high-performance computing.
Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
Strong programming skills in C++, Python, and familiarity with ML frameworks.
Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies.
Hands-on experience with Triton and/or JAX Pallas for custom kernel development.
Strong understanding of PTX, GPU ASM, and low-level GPU execution.
Extensive experience writing and optimizing custom GPU kernels in C++ and PTX.
Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks.
Experience working with large-scale LLM workloads (training or inference).
Nice-to-Haves
Experience with AMD GPUs and ROCm optimization.
Familiarity with JAX FFI and custom ML operator development.
Experience with efficient model serving frameworks (e.g., vLLM, TensorRT).
Experience with TPUs, XLA, or similar accelerator programming environments.
Contributions to open-source ML systems, compilers, or GPU kernels.
Benefits include
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Equal opportunity
Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
The full posting opens here — pay, setting and the full description, without leaving the list.