Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.
You'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
Engineering leadership experience building and scaling highly specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or closely related infrastructure.
Deep technical credibility at the inference layer, including hands-on understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs.
Ability to distinguish core inference-engine work from the routing, orchestration, and application layers above it, with opinions grounded in direct technical experience.
A strong record of recruiting, assessing, and retaining senior engineers, staff-level ICs, PhDs, and research-adjacent engineers in a production engineering environment.
Experience translating technically ambitious work into clear priorities, accountable ownership, execution plans, and durable engineering operating mechanisms.
Ability to remain close enough to the work to identify risks, pattern-match on difficult technical problems, and unblock teams without becoming a bottleneck or displacing technical ownership.
Preferred qualifications:
Experience leading teams responsible for LLM serving, vLLM, SGLang, model execution, inference performance, GPU kernels, compiler or runtime systems, or distributed AI infrastructure.
Experience scaling a small, senior-heavy engineering organization where the relevant talent market is narrow and technical quality matters more than headcount growth.
Experience integrating research-oriented or PhD talent into production teams, including setting expectations, structuring work, and building effective collaboration with product-focused engineers.
Strong judgment across organizational design, hiring, performance management, technical planning, execution cadence, and cross-functional decision-making.
Ability to represent the engineering organization credibly with open-source contributors, hardware partners, cloud providers, customers, candidates, and investors.
Bonus points if you have:
Built or led engineering teams working directly on GPU or accelerator-level inference performance, ML compilers, kernels, runtimes, or hardware-software co-design.
Contributed to or led teams around open-source ML systems projects such as vLLM, SGLang, PyTorch, Ray, Triton, XLA, ROCm, or related infrastructure.
Scaled an engineering organization through an inflection point while preserving high technical standards, fast iteration, and direct ownership.
Recruited successfully from a global, highly competitive ML systems talent pool and built relationships with technical communities beyond traditional candidate pipelines.
Led engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.
Logistics
Location: This role is based in San Francisco, California. Will consider relocation for exceptional candidates.
Compensation: Compensation will be determined based on background, skills, and experience. Offer will include a highly competitive base and meaningful equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
63 days ago
Member of Technical Staff, CI/CD Infrastructure
Inferact · San Francisco, California, United States
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.
Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.
Preferred qualifications:
Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.
Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.
Bonus points if you have:
Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.
Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
83 days ago
Member of Technical Staff, Inference
Inferact · San Francisco, California, United States
Member of Technical Staff, Inference
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Research & Engineering
Compensation
$200K – $400K • Offers Equity
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Deep understanding of transformer architectures and their variants.
Strong programming skills in Python with experience in PyTorch internals.
Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
Ability to read and implement model architectures and inference techniques from research papers.
Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
Preferred qualifications:
Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
Familiarity with RL frameworks and algorithms for LLMs.
Experience with multimodal inference (audio/image/video/text).
Contributions to open-source ML or system infrastructure projects.
Bonus points if you have:
Implemented core features in vLLM or other inference engine projects.
Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
Written widely-shared technical blogs or side projects on vLLM or LLM inference.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
84 days ago
Member of Technical Staff, TPU Performance Engineering
Inferact · San Francisco, California, United States
Member of Technical Staff, TPU Performance Engineering
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Research & Engineering
Compensation
$200K – $400K • Offers Equity
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.
Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.
Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.
Preferred qualifications:
Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.
Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.
Bonus points if you have:
Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.
Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
84 days ago
Member of Technical Staff, Developer Relations
Inferact · San Francisco, California, United States
Member of Technical Staff, Developer Relations
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Research & Engineering
Compensation
$200K – $400K • Offers Equity
Overview
Application
Overview
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference. This is not a generic DevRel role. We're looking for a inference systems educator-builder: someone who can understand vLLM as a deep LLM inference systems project, teach hard technical concepts clearly, and create public artifacts that help practitioners build better systems.
You'll write technical deep dives, build demos, create tutorials, contribute to docs and examples, host workshops, and help developers understand topics like KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems. Your work will shape how the broader AI infrastructure community learns, adopts, and builds with vLLM.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.
Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
Ability to credibly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.
Experience with vLLM or adjacent inference technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms, or similar systems.
A strong public portfolio of technical artifacts, such as blogs, tutorials, workshops, courses, OSS docs, benchmark posts, architecture explainers, conference talks, demos, or runnable repositories.
Ability to write and teach for practitioners without sounding like a content marketer.
Strong engineering judgment, product taste, and ability to turn raw technical material into useful developer education.
Preferred qualifications:
Prior work in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
Experience creating technical content that teaches reusable mental models, not just product features.
Experience contributing to developer-facing open source through docs, tutorials, examples, cookbooks, demos, or community support.
Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.
Ability to host workshops, create hands-on labs, present technical talks, and help developers move from concept to working code.
Bonus points if you have:
Written widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.
Built demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.
Contributed to open-source ML infrastructure, inference systems, developer tooling, or technical education projects.
Created practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.
Built a durable personal portfolio that demonstrates technical depth, taste, and a strong point of view.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
230 days ago
Member of Technical Staff, Kernel Engineering
Inferact · San Francisco, California, United States
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Strong systems programming skills in Rust, Go, or C++.
Experience designing and building high-performance distributed systems at scale.
Understanding of network protocols and high-performance I/O.
Ability to debug complex distributed systems issues.
Preferred qualifications:
Experience with ML serving infrastructure and disaggregated inference architecture.
Familiarity with GPU programming models and memory hierarchies.
Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics.
Track record of improving system reliability and performance at scale.
Bonus points if you have:
Prior experience in supporting large‑scale model training or inference environments.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
230 days ago
Member of Technical Staff, Cloud Orchestration
Inferact · San Francisco, California, United States
Member of Technical Staff, Cloud Orchestration
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Research & Engineering
Compensation
$200K – $400K • Offers Equity
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Strong experience with Kubernetes and container orchestration at scale.
Experience designing and implementing custom Kubernetes operators.
Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).
Experience managing GPU clusters and debugging hardware issues.
Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.
Preferred qualifications:
Experience with ML-specific orchestration tools (Ray, Slurm).
Knowledge of GPU scheduling, multi-tenancy, and resource optimization.
Familiarity with vLLM deployment patterns and configuration.
Track record of improving operational reliability for ML systems.
Bonus points if you have:
Experience deploying inference systems on large-scale GPU (1,000+) clusters.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
230 days ago
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)
Location
Remote
Employment Type
Full time
Location Type
Remote
Department
Research & Engineering
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.
You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.
Potential focus areas include:
Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures.
Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types.
Performance & Scale: Build the distributed systems that power inference at global scale—design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency.
Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring that enables teams worldwide to serve AI models without friction.
What We're Looking For
We're looking for engineers who thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.
Core Requirements:
Bachelor's degree or equivalent experience in computer science, engineering, or similar
Demonstrated ability to work autonomously and drive projects to completion without close supervision
Excellent asynchronous communication skills and ability to collaborate effectively across time zones
Strong track record of shipping high-impact work in complex technical environments
Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure
Technical Depth (strong in at least two):
CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture
High-performance distributed systems in Rust, Go, or C++
Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)
Kubernetes, container orchestration, and infrastructure-as-code at scale
Transformer architectures, KV-cache memory management, and model serving
Preferred Qualifications:
Contributions to vLLM or other major open-source ML/systems projects
Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)
Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies
Track record of improving system reliability and performance at scale
Written widely-shared technical blogs or impactful side projects in the ML infrastructure space
Logistics
Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.
Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.