Page 1
Loading more openings…
You've reached the end of the list.
Member of Technical Staff, Cloud Orchestration
Inferact · San Francisco, California, United States
Pay
$200k–400k
Setting
On-site
Member of Technical Staff, Cloud Orchestration
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Research & Engineering
Compensation
$200K – $400K • Offers Equity
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Strong experience with Kubernetes and container orchestration at scale.
Experience designing and implementing custom Kubernetes operators.
Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).
Experience managing GPU clusters and debugging hardware issues.
Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.
Preferred qualifications:
Experience with ML-specific orchestration tools (Ray, Slurm).
Knowledge of GPU scheduling, multi-tenancy, and resource optimization.
Familiarity with vLLM deployment patterns and configuration.
Track record of improving operational reliability for ML systems.
Bonus points if you have:
Experience deploying inference systems on large-scale GPU (1,000+) clusters.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Member of Technical Staff, Exceptional Generalist (Remote)
Inferact · US
Pay
$140k–250k
Setting
Remote
Member of Technical Staff, Exceptional Generalist (Remote)
Location
Remote
Employment Type
Full time
Location Type
Remote
Department
Research & Engineering
Overview
Application
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.
You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.
Potential focus areas include:
Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures.
Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types.
Performance & Scale: Build the distributed systems that power inference at global scale—design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency.
Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring that enables teams worldwide to serve AI models without friction.
What We're Looking For
We're looking for engineers who thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.
Core Requirements:
Bachelor's degree or equivalent experience in computer science, engineering, or similar
Demonstrated ability to work autonomously and drive projects to completion without close supervision
Excellent asynchronous communication skills and ability to collaborate effectively across time zones
Strong track record of shipping high-impact work in complex technical environments
Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure
Technical Depth (strong in at least two):
CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture
High-performance distributed systems in Rust, Go, or C++
Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)
Kubernetes, container orchestration, and infrastructure-as-code at scale
Transformer architectures, KV-cache memory management, and model serving
Preferred Qualifications:
Contributions to vLLM or other major open-source ML/systems projects
Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)
Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies
Track record of improving system reliability and performance at scale
Written widely-shared technical blogs or impactful side projects in the ML infrastructure space
Logistics
Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.
Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Inferact for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.