Page 1
Data Center Operations Systems Engineer (San Jose)
Lambda · San Jose, California, United States
$105k–140k
8 days ago
♡
Data Center Operations Systems Engineer (Atlanta)
Lambda · Atlanta, Georgia, United States
$89k–119k
12 days ago
♡
Senior Business Systems Architect
Lambda · San Francisco, California, United States
$206k–275k
15 days ago
♡
Senior Software Engineer – Billing & Identity
Lambda · San Francisco, California, United States
$380k–445k
15 days ago
♡
Engineering Manager, Fleet Engineering
Lambda · San Francisco, California, United States
$297k–440k
18 days ago
♡
Technical Success Engineer
Lambda · San Francisco, California, United States
$251k–335k
20 days ago
♡
Senior Site Reliability Engineer - Fleet
Lambda · San Francisco, California, United States
$240k–356k
25 days ago
♡
Senior Software Engineer - Managed Kubernetes
Lambda · San Francisco, California, United States
$266k–395k
26 days ago
♡
Loading more openings…
You've reached the end of the list.
Data Center Operations Systems Engineer (San Jose)
Lambda · San Jose, California, United States
Pay
$105k–140k
Setting
On-site
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Jose Data Centers 5 days; shift work
What You'll Do
- Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured
- Troubleshoot hardware and software issues in some of the world’s most advanced systems
- Document data center layout and network topology in DCIM software
- Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments
- Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers
- Work closely with HW Support team to ensure data center infrastructure-related support tickets are resolved
- Work with RMA team to ensure faulty parts are returned and replacements are ordered
- Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers
You
- Are familiar with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Are Fluent in English and Spanish
- Are someone who pays attention to detail and has the ability to follow instructions
- Are action-oriented and have a strong willingness to learn
- Are willing to travel for bring up of new data center locations
Nice to Have
- Experience with troubleshooting server hardware
- Experience with/or knowledge of network topology
- Familiarity with ticketing systems like JIRA and Zendesk
- Experience with Linux administration
- Experience with working in large-scale distributed data center environments
- Experience with Supermicro & Nvidia hardware
Salary Range Information
This is a salaried non-exempt role, eligible for overtime. The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Data Center Operations Systems Engineer (Atlanta)
Lambda · Atlanta, Georgia, United States
Pay
$89k–119k
Setting
On-site
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our Atlanta, GA Data Centers 5 days; shift work
What You'll Do
- Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured
- Troubleshoot hardware and software issues in some of the world’s most advanced systems
- Document data center layout and network topology in DCIM software
- Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments
- Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers
- Work closely with HW Support team to ensure data center infrastructure-related support tickets are resolved
- Work with RMA team to ensure faulty parts are returned and replacements are ordered
- Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers
You
- Are familiar with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Are someone who pays attention to detail and has the ability to follow instructions
- Are action-oriented and have a strong willingness to learn
- Are willing to travel for bring up of new data center locations
Nice to Have
- Experience with troubleshooting server hardware
- Experience with/or knowledge of network topology
- Familiarity with ticketing systems like JIRA and Zendesk
- Experience with Linux administration
- Experience with working in large-scale distributed data center environments
- Experience with Supermicro & Nvidia hardware
Salary Range Information
This is a salaried non-exempt role, eligible for overtime. The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Business Systems Architect
Lambda · San Francisco, California, United States
Pay
$206k–275k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About the Role
We are seeking an experienced Senior Business Systems Architect to lead and innovate across our enterprise systems landscape. This role will be pivotal in driving operational excellence through the design, implementation, and optimization of scalable business processes and systems integrations. You will collaborate with cross-functional teams to create seamless workflows across financial systems, supply chain, and customer-facing platforms while leveraging process automation tools and middleware solutions.
Your responsibilities will include architecting end-to-end solutions, managing complex system integrations, and building a strategic roadmap for enterprise system enhancements. This role requires strong leadership, deep technical expertise, and a proactive approach to solving business challenges.
What You’ll Do
- Strategic Leadership: Collaborate with stakeholders across functions to develop and execute scalable system roadmaps, aligning business goals with system capabilities.
- Solution Design & Implementation: Architect and deploy advanced workflows, configurations, and modules in enterprise platforms like NetSuite, Coupa, Stripe, and Chargebee.
- Tooling Selection & Standardization: Lead the evaluation, selection, and implementation of tools and systems across the business to ensure alignment with operational needs, scalability, and long-term goals. Establish standards for system usage and integration.
- Integration & Automation: Manage integrations and optimize iPaaS tools (e.g., Workato, Mulesoft) to ensure seamless data flow, interoperability, and automated processes across platforms.
- Compliance & Optimization: Ensure SOX compliance through robust internal controls, system audits, and governance while driving continuous system improvements.
- Support & Analytics: Troubleshoot complex issues, mentor team members, and deliver actionable insights through advanced reporting and dashboards to support decision-making.
You
- Experience & Certification:
- 10+ years of experience in business systems architecture or administration, including hands-on implementation of enterprise systems.
- Strong understanding of Finance technology stacks, including NetSuite, Coupa, Supply Chain, Asset Management, Billing systems and related business applications
- Certifications in relevant platforms (e.g., NetSuite Administrator, iPaaS) strongly preferred.
- Technical Expertise:
- Proficiency in languages like JavaScript/SQL/ HTML/CSS/Python.
- Strong experience with middleware and low/no-code platforms (e.g., Workato, n8n, Mulesoft, Alteryx) for system integrations & workflow automations.
- Process Automation:
- Demonstrated success in deploying process automation tools to streamline operations and increase efficiency.
- Compliance & Controls:
- Strong knowledge of ITGCs (IT General Controls) and SOX compliance.
- Leadership & Communication:
- Excellent communication skills to effectively collaborate across all levels of leadership from technical staff to senior executives.
- Proven ability to present complex technical solutions in a clear and compelling manner
- AI Automation & Development
- Hands-on experience leveraging AI and machine learning tools to design, build, and deploy intelligent automation solutions that enhance business processes.
- Proficiency with AI-powered platforms and frameworks (e.g., OpenAI, Anthropic, LangChain, or similar) to develop custom workflows, copilots, and agentic solutions.
- Proven ability to identify high-impact automation opportunities and translate them into scalable AI-driven implementations across finance and enterprise systems.
Nice to Have
- Proficiency in system customization tools like SuiteScript, Salesforce Flow, or equivalent.
- Experience with building Internal AI apps on AWS.
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Software Engineer – Billing & Identity
Lambda · San Francisco, California, United States
Pay
$380k–445k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
What You’ll Do
- Architect, design, and develop complex distributed systems and computer applications for processing massive amounts of data and usage events
- Design and build report interfaces and data feeds for downstream financial systems to ensure financial accuracy.
- Create consumer software products for billing and pricing features using Golang and Python programming languages by iteratively taking technical design requirements from the design team and product requirements from the product teams.
- Responsible for improving software quality through thoughtful code reviews, appropriate software testing, proper rollout, monitoring, and proactive changes.
- Establish best practices for enabling continuous integrations deployments a week and setup a clean process to rollback changes as needed.
- Track system and software key performance indicators (KPIs), revenue impact, response times, and latency to make data-driven recommendations for continuous improvement.
Some telecommuting permitted (hybrid).
You
- Master’s degree in Computer Science or a related field plus 5 years of experience as a Software Engineer or related occupation.
- Must have at least 1 year of prior work experience in each of the following:
1. Industry-standard coding languages, including Python and GoLang
2. Databases, Storage & Analysis, including SQL, Hive, Presto, DynamoDB, and S3
3. Software development tools and revision control systems, including VIM, GIT, and shell scripting
4. Building highly-scalable performant solutions
5. Applying algorithms and core computer science concepts to real world systems as evidenced by recognizing and matching patterns from different areas of computer science in production systems
6. Designing scalable distributed systems with established partition tolerance, consistency, and availability guarantees.
#LI-DNI #LI-NDI #LI-DNP
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering Manager, Fleet Engineering
Lambda · San Francisco, California, United States
Pay
$297k–440k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About the Role
Fleet Engineering owns the full lifecycle of Lambda's production systems infrastructure — new product introduction, deployment, operation, and reliability of our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The Fleet Engineering teams:
- HPC Deployments — Turns bare metal into production-ready capacity: ensure firmware leveling, system burn-in to shake out early failures, and validation of server performance and correctness, through to the InfiniBand fabric and GPU clusters.
- Fleet Reliability — Day-2 operations across the fleet. Keeps systems healthy and keeps as much of the fleet in service for as much of its useful life as possible.
- Fleet Orchestration / Data — Owns our production source of truth system. Synchronizes data from upstream systems and holds the line on correctness and quality, because everything automated downstream depends on it.
- Fleet Orchestration / Automation — Owns the workflow orchestration system which people use to safely work on fleet systems for workflows that include: locking hosts, running firmware leveling jobs, OS installs, burn-in and validation, and reporting on work in flight and its results.
- Fleet Foundation — Builds the host enablement tooling systems: OS and ZTP switch provisioning, firmware management, out-of-band access, and power management.
The work is highly cross-functional, carries executive visibility, and has a direct impact on Lambda and our customers. Fleet Engineering is at the forefront of delivering on-time, high-quality GPU capacity while driving efficiency at scale.
We are hiring multiple Engineering Managers for the following teams: Fleet Reliability, HPC Deployments, Fleet Foundation, Fleet Orchestration / Automation. This is a single application for all of them: you apply once, we get to know you, and we match you to the team where your strengths land best.
We value diverse backgrounds, experiences, and skills, and we're excited to hear from candidates who bring a unique perspective. If you don't exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role. Your application is not a waste of our time.
What You'll Do
- Lead and grow a distributed team of top-talent engineers responsible for the deployment and operation of production systems infrastructure.
- Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders.
- Identify opportunities for efficiency gains in the tools, processes, and automation that teams across the organization rely on day to day.
- Give stakeholders clear visibility into project progress, risks, and outcomes.
- Participate in qualification efforts for new technologies entering our production deployments.
- Drive outcomes by managing staff allocation, project priorities, deadlines, and deliverables.
- Hold regular 1:1s, give constructive feedback, and support career development for your team.
- Contribute to reliability through participation in our Incident Management and Review programs.
You
- Have 3+ years leading or managing engineers, in AI/ML infrastructure or another large-scale compute environment.
- Have owned production systems with real SLAs, and can balance keeping things running against long-term, high-impact work — paying down toil and technical debt along the way.
- Work confidently in Linux and can debug across the OS, hardware, and networking layers.
- Can lead technical design on medium-to-large efforts: take an ambiguous problem, write the doc, drive alignment across teams, and ship.
- Work well under deadlines and structured project plans, and can tactfully negotiate changes to timelines when reality demands it.
- Collaborate effectively with peer engineering managers on efforts that cut across deployment and operations.
- Build high-performing teams deliberately — through hiring, upskilling, planned skills redundancy, performance management, and clear expectations.
- Have excellent problem-solving and troubleshooting instincts.
- Are excited about working at the intersection of hardware, software, and physical datacenter builds.
- Leave systems, and the teammates around you, better than you found them.
Nice to Have
Depth in any one of these is a strong signal, and helps us place you on the right team. Nobody has all of them.
- Linux systems administration, TCP/IP networking, automation, and scripting.
- Bare metal provisioning and lifecycle management — PXE, Redfish, IPMI, BMC, DHCP, DNS.
- Strong coding ability in at least one language, plus comfort with APIs, distributed systems, and automation pipelines.
- The technologies underpinning our cloud business: GPU acceleration, virtualization, cloud computing.
- Datacenter physical infrastructure: racks, switches, InfiniBand fabric, power domains.
- Network source-of-truth or DCIM tooling (NetBox or similar), and data quality practice at scale.
- Building Linux distributions, or managing OS customization and imaging.
- Incorporating AI-assisted development tools into engineering workflows — code generation, debugging, test development, documentation.
- Customer awareness, empathy, and diplomacy.
- Bachelor's degree or equivalent experience in a technical field.
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Technical Success Engineer
Lambda · San Francisco, California, United States
Pay
$251k–335k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About this role
The Superintelligence Technical Success Engineer is part of Lambda's Superintelligence business unit, dedicated to our largest, most strategic customers operating in the most complex environments. This role owns taking signed deployment from contract to a live, fully operational production environment. The role requires strong technical acumen to validate the build against what was promised, spot gaps, and work effectively with engineering and infrastructure teams to get issues resolved and requirements clearly understood.
You'll work alongside the account team, to make sure the customer's requirements and expectations are clearly understood and addressed throughout deployment. This role calls for someone who's ready to dive in wherever the engagement needs them.
What You'll Do
- Take signed deployments, from contract, to live, to working production environments, validating configuration, connectivity, storage, and compute against what was promised
- Bring technical depth to validation and troubleshooting conversations — asking the right questions, spotting gaps, and working closely with engineering and infrastructure teams to drive issues to resolution
- Coordinate with Infrastructure, Engineering, Product, and Data Center teams to close technical dependencies and resolve blockers
- Own a current, accurate technical picture of the deployment — what's built, what's open, what's at risk
- Keep stakeholders informed with clear, regular status updates: RAG status, top risks, and what's being done about them
- Guide the customer through onboarding to their first successful production workload
- Be the customer's go-to technical contact through deployment and early production
- Feed recurring technical patterns back into reusable runbooks, checklists, or automation
- Transition out once the customer is stable and self-sufficient, keeping the broader account team informed along the way
You
- 4+ years of hands-on technical experience with GPU/HPC infrastructure, cloud platforms, Kubernetes, or large-scale Linux systems
- Comfortable being the technical voice in the room able to validate builds, spot gaps, and hold engineering teams accountable for resolving them
- Strong troubleshooting instincts and a willingness to get into the weeds of networking, storage, or compute issues
- Track record of coordinating across engineering and infrastructure teams to close out technical dependencies
- Clear written and verbal communication for status updates and technical documentation
- A bias toward diving in and taking ownership, rather than waiting for a fully defined process
Nice to Have
- Exposure to large-scale GPU cluster deployments
- Familiarity with project/program tracking tools and structured status reporting
- Experience building runbooks or checklists that outlived the engagement they were built for
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Site Reliability Engineer - Fleet
Lambda · San Francisco, California, United States
Pay
$240k–356k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
What You’ll Do
- Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively
- Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible
- Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand
- Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely
- Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teams
- Participate in on-call rotations and lead incident response for cluster-level problems
- Contribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiency
You
- 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
- Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
- Strong understanding of Linux-based systems in a distributed environment
- Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
- Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.
- Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
- Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
- Have excellent problem-solving and troubleshooting skills and an innate attention to detail
- Passion for continuous improvement and innovation
Nice to Have
- Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
- Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
- Experience building and/or operating HPC resources.
- Depth in the NVIDIA hardware and firmware ecosystem
- Experience with data center power and thermal design
- Background in chaos engineering or similar reliability testing methodologies
- Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Software Engineer - Managed Kubernetes
Lambda · San Francisco, California, United States
Pay
$266k–395k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About the Role
We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in shaping the architecture, reliability, and automation of our Kubernetes-based infrastructure, which powers mission-critical workloads across our global platform.
Lambda is building the AI Cloud of the future. We are seeking a Senior Software Engineer to help our development of our Managed Kubernetes platform. Think GKE, but purpose-built for AI workloads and running on bare metal. In this role, you will help build the infrastructure that powers the next generation of AI training and inference at scale.
As a Senior Engineer on our Orchestration team, you will contribute to Lambda's managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps. You'll work at the intersection of distributed systems, GPU-accelerated computing, and Cloud Native infrastructure to build systems that are reliable, performant, and elegantly simple for our customers.
This is not a role for someone who just operates Kubernetes; it's a role for an engineer who understands how compute, network, storage, and security interact, and can build solutions that account for that context — even while focused primarily on the orchestration layer. You'll be working closely with NVIDIA's open-source ecosystem, and partnering with internal teams across the stack to deliver a world-class managed platform.
What You’ll Do
- Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management — provisioning, upgrades, patching, and deletion
- Build GPU-aware orchestration systems, working within the platform architecture to support GPU scheduling and resource allocation
- Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect
- Write resilient systems that handle failure gracefully — timeouts, retries, backoff, and degraded-mode operation — across large-scale distributed environments
- Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, and multi-model deployment patterns
- Build internal tools and CLIs that let ML/AI teams deploy and monitor their own inference services
- Support and debug production issues through on-call rotation
Required Qualifications
- Have 6+ years of experience in software engineering, with a track record of owning significant technical scope within a team (e.g., driving a project from design through production, or acting as a de facto tech lead on a workstream)
- Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
- Solid grasp of distributed systems fundamentals — fault tolerance, graceful degradation, and failure handling in large-scale environments
- Experience operating the control plane and low-level pieces of large-scale Kubernetes clusters
- Experience with observability at scale: Prometheus, Grafana, distributed tracing, and building actionable alerting systems
- Strong programming skills in Go and Python; ability to collaborate effectively on shared codebases
- Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
- Take pride in owning and delivering core components of products and platforms
Preferred Qualifications
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS, or similar) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar
- Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
- Exposure to storage architecture for AI/ML workloads
- Past contributions to CNCF projects or Kubernetes SIGs a plus
If you don’t meet all of these requirements but believe you may be a good fit, please still apply and provide a cover letter that helps us understand your experience and readiness for this role.
Why Lambda
Lambda is building the essential infrastructure for the AI era. We're not just another cloud provider: we're a company founded by ML practitioners, for ML practitioners. Our customers include leading AI research labs and enterprises pushing the boundaries of what's possible with artificial intelligence.
What makes this role special:
- You'll be building core platform services the world's largest AI companies will consume
- NVIDIA partnership: Deep integration with NVIDIA's GPU and networking stack, working with cutting-edge open-source tooling
- Real technical challenges: Massive scale GPU clusters and the unique demands of AI workloads
- Cross-stack exposure: Work at the intersection of Kubernetes, networking, storage, and compute — gaining depth across the full infrastructure stack that powers AI workloads, not just the orchestration layer
- Direct impact: Your work enables AI breakthroughs. Every model trained on Lambda benefits from systems you build
- World-class team: Work alongside engineers with deep expertise in ML, systems, and infrastructure
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Site Reliability Engineer - Managed Kubernetes
Lambda · San Francisco, California, United States
Pay
$240k–356k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.
What You’ll Do
- Operate and maintain bare-metal Kubernetes clusters, scaling up to thousands of nodes
- Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
- Participate in a well-managed on-call rotation for critical incidents
- Assist customers with Kubernetes questions, workload integration, storage, and authentication
- Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues
- Use Python and Golang to create tooling and automate the validation of platform quality.
- Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes
- Develop automation for cluster lifecycle management: provisioning, upgrades, patching, and deletion.
- Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.
You
- 6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and systems
- Strong programming skills in Go and Python; experience with GitOps (e.g., ArgoCD), Helm, and Kubernetes operators
- Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or similar)
- Can work either independently with limited direction or as part of a team
- Can work with customers during incidents either via tickets, live messaging, or as part of a larger call.
- Familiarity with observability tools like Prometheus, Grafana, FluentBit, and CI/CD pipelines
- Proven experience provisioning Kubernetes using tools such as kubeadm, Cluster API, or similar
Nice To Have
- Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience
- Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters
- Hybrid or multi-cloud Kubernetes environment experience
- Contributions to CNCF projects or Kubernetes SIGs
Why Join Us
- Work on cutting-edge Managed Kubernetes platforms for AI/ML workloads
- Influence the platform roadmap and help shape operations and reliability best practices
- Collaborate with a highly skilled engineer
- Opportunity to mentor and grow within a fast-growing, technology-driven environment
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Staff Software Engineer - Compute
Lambda · Bellevue, Washington, United States
Pay
$314k–465k
Setting
Hybrid
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About the Role
As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.
The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.
What You'll Do
We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:
- Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
- Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
- Guide design of compute platform multi-tenant security model
- Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
- Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
- Work with customers on translating vague customer technical requirements into concrete engineering deliverables.
- Set engineering standards and lead design reviews for mission-critical cloud software at scale.
Who You are
- 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
- Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
- Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
- Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
- Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
- Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.
Nice to Have
- Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
- Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
- Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
- Experience with Cloud Service Provider Kubernetes offerings.
- Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Listed by Lambda for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.