Engineering Manager, Platform Infrastructure (Foundations)
Location
San Francisco
Employment Type
Full time
Department
Engineering
Compensation
Estimated base salary across a range of levels: $270K – $320K • Offers Equity
Compensation:
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted accordingly.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Flexible Time Off
Commute reimbursement
100% of in-office meals covered
Overview
Application
Anyscale Platform Engineering Leader
About Anyscale:
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role:
Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams.
Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless.
In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges.
You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components.
You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful.
We'd love to hear from you if you have:
Solid engineering management experience leading productive, building high-performing teams
Record of helping teams scale quickly while maintaining a good culture
Ability to ensure a high hiring bar, motivate team, coach/mentor, and handle performance management issues.
Deep technical knowledge and experience in distributed systems and prior experience working on kubernetes, VMs.
A great track record of execution
A sense of urgency, a mindset towards achieving results, and excellent prioritization skills
Effective communication
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
Anyscale is looking for a Software Engineer to join the Platform and Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the team, we build the scalable, secure, and robust backbone that enables this vision, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world.
Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads.
We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale's cloud platform.
You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers.
A snapshot of projects you may work on
Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
Optimize control plane components for large-scale, distributed AI/ML workloads
Build intelligent scheduling and resource management systems for heterogeneous compute clusters
Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
Support and optimize accelerator integration (e.g., GPUs, TPUs).
Handle container image management and dependency resolution for distributed workloads
Participate in code reviews, design and architecture discussions
Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
We'd love to hear from you if have
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
3+ years of experience writing high-quality production code
Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
Deep understanding of networking, security, and authentication mechanisms in cloud environment
Familiarity with observability stacks (Prometheus, Grafana etc)
Proficiency in Go and Python
Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
As a Forward Deployed Engineer at Anyscale, you will partner directly with our most strategic customers, including Spanish-speaking customers across Latin America and other regions, to ensure they achieve meaningful business outcomes with Ray and the Anyscale platform. Embedded within customer teams, you’ll act as a trusted advisor, aligning technical solutions with customer priorities, accelerating time-to-value, and driving adoption at scale.
You’ll work across customer organizations — from technical leadership to individual contributors — to scope and deliver impactful solutions. By connecting insights from the field back to our product and engineering teams, you’ll help shape Anyscale’s roadmap and ensure we remain focused on solving our customers’ most critical challenges.
In this role, you will:
Work onsite with key customers to lead proof-of-value engagements, deployments, and enterprise adoption
Translate business objectives into technical solutions that demonstrate clear ROI and strategic impact
Build and deliver high-impact demos, reference architectures, and enablement programs tailored to customer needs
Act as a trusted advisor across all levels of the organization, ensuring confidence in Anyscale and Ray
Collaborate closely with sales, product, and engineering to accelerate deals, unblock challenges, and drive long-term success
Provide structured feedback from customer engagements to influence product direction and go-to-market strategy
We’d love to hear from you if you have:
Fluency in Spanish and English is required, with the ability to work directly with Spanish-speaking technical and executive stakeholders
5+ years of customer-facing experience in forward deployed engineering, solutions architecture, field engineering, or software engineering
Strong technical foundation with Ray, or the demonstrated ability to ramp quickly and apply Ray to real-world use cases
Hands-on experience with ML training and inference workloads, including distributed training, model serving, and the performance and cost tradeoffs involved
Working knowledge of Kubernetes and container orchestration, including deploying and operating workloads in Kubernetes-based environments
Proven success driving enterprise adoption of complex SaaS, infrastructure, or ML/AI solutions
Experience engaging both executive and technical stakeholders, tailoring communication to each audience
A customer-first mindset with a track record of delivering measurable business impact
Willingness to travel frequently and spend extended time embedded with customers
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Engineering
84 days ago
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Location
San Francisco; Palo Alto
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Compensation
Target Base Salary:
$215K – $275K • Offers Equity
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 95%
401k Retirement Plan
Wellness & Education Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
In the office? Lunch is on us!
Overview
Application
About Anyscale:
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role:
Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision.
Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads.
We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform.
You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers.
A snapshot of projects you may work on
Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
Optimize control plane components for large-scale, distributed AI/ML workloads
Build intelligent scheduling and resource management systems for heterogeneous compute clusters
Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
Support and optimize accelerator integration (e.g., GPUs, TPUs).
Handle container image management and dependency resolution for distributed workloads
Participate in code reviews, design and architecture discussions
Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
We'd love to hear from you if have
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
3+ years of experience writing high-quality production code
Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
Deep understanding of networking, security, and authentication mechanisms in cloud environment
Familiarity with observability stacks (Prometheus, Grafana etc)
Proficiency in Go and Python
Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Distributed LLM Inference Engineer
Location
San Francisco; Palo Alto
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Compensation
Target Base Salary:
$170K – $245K • Offers Equity
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 99%
401k Retirement Plan
Wellness & Education Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in office meals covered
Overview
Application
About Anyscale
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.
As part of this role, you will
Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices
We'd love to hear from you if you have
Familiarity with running ML inference at large scale with high throughput and low latency
Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
Solid understanding of distributed systems, ML inference challenges
Bonus points!
ML Systems knowledge
Experience using Ray
Work closely with community on LLM engines like vLLM, TensorRT-LLM
Contributions to deep learning frameworks (PyTorch, TensorFlow)
Contributions to deep learning compilers (Triton, TVM, MLIR)
Prior experience working on GPUs / CUDA
Compensation
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in-office meals covered
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Software Engineer, Ray Core
Location
Bengaluru, Karnataka
Employment Type
Full time
Department
Engineering
Overview
Application
About Anyscale:
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend.
About the Ray Core Team
The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray.
A snapshot of projects you can work on:
- Optimizing performance of large-scale workloads on Ray
- Stability and stress testing infrastructure
- Improving fault tolerance (HA)
As part of this role, you will:
Develop high quality open source software to simplify distributed programming (Ray)
Identify, implement, and evaluate architectural improvements to Ray core
Improve the testing process for Ray to make releases as smooth as possible
Communicate your work to a broader audience through talks, tutorials, and blog posts
We'd love to hear from you if have:
At least 2 year of relevant work experience
Solid background in algorithms, data structures, system design
Experience in building scalable and fault-tolerant distributed systems
Knowledge of distributed model training and inference (e.g. tensor parallel, pipeline parallel) is preferred
Knowledge of GPU programming is preferred
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.