Founding Product Designer
Head of Engineering
Overview
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.
You'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
Engineering leadership experience building and scaling highly specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or closely related infrastructure.
Deep technical credibility at the inference layer, including hands-on understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs.
Ability to distinguish core inference-engine work from the routing, orchestration, and application layers above it, with opinions grounded in direct technical experience.
A strong record of recruiting, assessing, and retaining senior engineers, staff-level ICs, PhDs, and research-adjacent engineers in a production engineering environment.
Experience translating technically ambitious work into clear priorities, accountable ownership, execution plans, and durable engineering operating mechanisms.
Ability to remain close enough to the work to identify risks, pattern-match on difficult technical problems, and unblock teams without becoming a bottleneck or displacing technical ownership.
Preferred qualifications:
Experience leading teams responsible for LLM serving, vLLM, SGLang, model execution, inference performance, GPU kernels, compiler or runtime systems, or distributed AI infrastructure.
Experience scaling a small, senior-heavy engineering organization where the relevant talent market is narrow and technical quality matters more than headcount growth.
Experience integrating research-oriented or PhD talent into production teams, including setting expectations, structuring work, and building effective collaboration with product-focused engineers.
Strong judgment across organizational design, hiring, performance management, technical planning, execution cadence, and cross-functional decision-making.
Ability to represent the engineering organization credibly with open-source contributors, hardware partners, cloud providers, customers, candidates, and investors.
Bonus points if you have:
Built or led engineering teams working directly on GPU or accelerator-level inference performance, ML compilers, kernels, runtimes, or hardware-software co-design.
Contributed to or led teams around open-source ML systems projects such as vLLM, SGLang, PyTorch, Ray, Triton, XLA, ROCm, or related infrastructure.
Scaled an engineering organization through an inflection point while preserving high technical standards, fast iteration, and direct ownership.
Recruited successfully from a global, highly competitive ML systems talent pool and built relationships with technical communities beyond traditional candidate pipelines.
Led engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.
Logistics
Location: This role is based in San Francisco, California. Will consider relocation for exceptional candidates.
Compensation: Compensation will be determined based on background, skills, and experience. Offer will include a highly competitive base and meaningful equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Product Marketing Manager
Overview
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We’re looking for a Product Marketing Manager to accelerate vLLM’s position as the leading open-source inference engine and publicize Inferact’s role in driving its development. This role sits at the intersection of product marketing, events, partner marketing, community, and growth. You’ll help turn a high volume of conference and ecosystem demand into a proactive GTM motion that builds brand awareness, community engagement, and qualified momentum around vLLM and Inferact
You’ll own end-to-end execution for conferences, partner events, hosted events, landing pages, speaker coordination, booth needs, collateral, follow-up workflows, and the many details that make technical events successful. You’ll develop and strengthen relationships between Inferact and major corporate and ecosystem partners, while shaping the messaging and presence of both vLLM and Inferact, while developing . You’ll work closely with product, engineering, leadership, partners, and design collaborators. This is a hands-on 0-to-1 role for someone with strong taste, high urgency, and an AI-native way of working.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in marketing, product, computer science, business, communications, or a related field.
Strong product marketing or GTM fundamentals, including positioning, messaging, audience definition, narrative development, and practical execution.
AI-native operating style, with regular use of modern AI tools to accelerate writing, research, planning, landing pages, workflows, asset drafts, and execution.
Experience owning events, conferences, partner activations, launches, or community programs end-to-end, including logistics, timelines, vendors, speakers, assets, landing pages, and follow-up.
Strong taste and judgment around brand, tone of voice, event experience, marketing quality, and developer-facing credibility.
Ability to operate autonomously in a startup environment, proactively recommend what should happen next, and execute without needing every step prescribed.
Ability to work credibly with product, engineering, founders, partners, design agencies, and technical communities.
Preferred qualifications:
Experience marketing to developers, infrastructure engineers, ML engineers, technical founders, open-source communities, or AI / developer tools audiences.
Experience with AI infrastructure, developer tools, model serving, cloud platforms, open-source infrastructure, or other technical products.
Experience running partner marketing or co-marketing programs with ecosystem partners such as cloud providers, hardware companies, infrastructure platforms, or developer communities.
Ability to personally create or coordinate landing pages, event pages, simple websites, forms, briefs, and campaign assets using AI, no-code tools, or lightweight technical workflows.
Experience building 0-to-1 marketing programs, event playbooks, launch motions, or community programs in a startup or high-ambiguity environment.
Bonus points if you have:
Built an event or community strategy that materially improved brand awareness, partner momentum, developer mindshare, or qualified pipeline.
Worked in or around AI infrastructure, open-source infrastructure, developer tools, data infrastructure, cloud, GPU, or technical SaaS ecosystems.
Created high-quality technical marketing assets, launch narratives, event experiences, partner campaigns, or developer-facing content that you can walk through in detail.
Partnered with design agencies or external creative teams while maintaining speed, quality, and brand consistency.
Helped a startup move from reactive marketing execution to a proactive 6-12 month GTM, events, or community plan.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Member of Technical Staff, CI/CD Infrastructure
Member of Technical Staff, AMD GPU Performance Engineering
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.
Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.
Preferred qualifications:
Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.
Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.
Bonus points if you have:
Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.
Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Member of Technical Staff, Inference
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, Developer Relations
Member of Technical Staff, Kernel Engineering
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
Proficiency in C++ and Python with demonstrated ability to write high-performance code.
Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.
Obsession with benchmarks and squeezing every percentage point of speedup.
Preferred qualifications:
Experience with ML-specific kernel optimization (FlashAttention, fused kernels).
Knowledge of quantization techniques (INT8, FP8, mixed-precision).
Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).
Experience with compiler technologies (LLVM, MLIR, XLA).
Bonus points if you have:
Kernel-related contributions to vLLM or other inference engine projects.
Contributions to open-source GPU, ML systems, or compiler optimization projects
Written deep technical blogs on GPU optimization.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Member of Technical Staff, Performance and Scale
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Strong systems programming skills in Rust, Go, or C++.
Experience designing and building high-performance distributed systems at scale.
Understanding of network protocols and high-performance I/O.
Ability to debug complex distributed systems issues.
Preferred qualifications:
Experience with ML serving infrastructure and disaggregated inference architecture.
Familiarity with GPU programming models and memory hierarchies.
Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics.
Track record of improving system reliability and performance at scale.
Bonus points if you have:
Prior experience in supporting large‑scale model training or inference environments.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
The full posting opens here — pay, setting and the full description, without leaving the list.