At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
Anyscale is looking for a Software Engineer to join the Platform and Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the team, we build the scalable, secure, and robust backbone that enables this vision, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world.
Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads.
We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale's cloud platform.
You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers.
A snapshot of projects you may work on
Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
Optimize control plane components for large-scale, distributed AI/ML workloads
Build intelligent scheduling and resource management systems for heterogeneous compute clusters
Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
Support and optimize accelerator integration (e.g., GPUs, TPUs).
Handle container image management and dependency resolution for distributed workloads
Participate in code reviews, design and architecture discussions
Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
We'd love to hear from you if have
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
3+ years of experience writing high-quality production code
Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
Deep understanding of networking, security, and authentication mechanisms in cloud environment
Familiarity with observability stacks (Prometheus, Grafana etc)
Proficiency in Go and Python
Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Location
San Francisco; Palo Alto
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Compensation
Target Base Salary:
$215K – $275K • Offers Equity
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 95%
401k Retirement Plan
Wellness & Education Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
In the office? Lunch is on us!
Overview
Application
About Anyscale:
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role:
Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision.
Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads.
We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform.
You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers.
A snapshot of projects you may work on
Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
Optimize control plane components for large-scale, distributed AI/ML workloads
Build intelligent scheduling and resource management systems for heterogeneous compute clusters
Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
Support and optimize accelerator integration (e.g., GPUs, TPUs).
Handle container image management and dependency resolution for distributed workloads
Participate in code reviews, design and architecture discussions
Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
We'd love to hear from you if have
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
3+ years of experience writing high-quality production code
Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
Deep understanding of networking, security, and authentication mechanisms in cloud environment
Familiarity with observability stacks (Prometheus, Grafana etc)
Proficiency in Go and Python
Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Machine Learning Engineer, Customer Engineering
Location
San Francisco
Employment Type
Full time
Location Type
Hybrid
Department
Customer Solutions Group
Compensation
$170K – $199K • Offers Equity
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 95%
401k Retirement Plan
Wellness & Education Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
In the office? Lunch is on us!
Overview
Application
About Anyscale
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
The Customer Engineer will play a crucial role in the customers’ post-sale journey - helping them to onboard, adopt and grow on Anyscale, troubleshooting and resolving open customer tickets and driving consumption.
Anyscale is an ever evolving platform and hence will require close co-ordination with our engineering teams to debug complex issues. This is an exciting role for those who are technically curious and passionate about ML/AI, LLM, vLLM and the role of AI in next generation applications. It’s an opportunity to make a significant impact in a collaborative, fast-paced environment while building a new segment in this space.
In this role, you’ll be able to
Resolve customer issues and help in their successful adoption of Anyscale platform
Be a technical advisor, and internal champion for our key customers
Own customer issues end-to-end, from troubleshooting, triaging, escalations and eventual resolution
Participate in our follow-the-sun customer support model to ensure continuity in resolving high priority tickets
Keep track of open customer bugs and feature requests to influence prioritization and provide timely customer updates upon resolution
Contribute towards improvement of internal tools and documentation of playbooks, guides and best practices etc. based on observed patterns
Habitually provide feedback and collaborate cross-functionally with product and engineering teams to address customer issues with a focus on improving the product experience
Build and maintain strong relationships with technical stakeholders within customer accounts
Qualifications
7+ years of experience in a Machine Learning role in a dynamic, fast-paced, startup-like environment
Strong organizational skills and ability to manage multiple customer needs simultaneously
Proficient at developing data pipelines for training, fine-tuning and inference/serving of LLMs
Experience running and optimizing infrastructure for distributed ML workloads on the major cloud platforms (AWS/EKS, GCP/GKE or Azure/AKS)
Excellent communication and interpersonal skills.
Strong sense of ownership, self-motivation and eagerness to acquire new skills and do new things
Willingness to up-level the knowledge and skills of your peers through mentorship, trainings and shadowing
Bonus
Experience with Ray
Knowledge of MLOps platforms
Knowledge of container orchestration platforms (e.g., Kubernetes), infrastructure as code (e.g. Terraform), CI/CD tools (e.g. Github Actions)
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review theNotice of E-Verify Participation and theRight to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.
Distributed LLM Inference Engineer
Location
San Francisco; Palo Alto
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Compensation
Target Base Salary:
$170K – $245K • Offers Equity
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 99%
401k Retirement Plan
Wellness & Education Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in office meals covered
Overview
Application
About Anyscale
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.
As part of this role, you will
Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices
We'd love to hear from you if you have
Familiarity with running ML inference at large scale with high throughput and low latency
Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
Solid understanding of distributed systems, ML inference challenges
Bonus points!
ML Systems knowledge
Experience using Ray
Work closely with community on LLM engines like vLLM, TensorRT-LLM
Contributions to deep learning frameworks (PyTorch, TensorFlow)
Contributions to deep learning compilers (Triton, TVM, MLIR)
Prior experience working on GPUs / CUDA
Compensation
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
Stock Options
Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in-office meals covered
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by AnyScale for a position based in the United States. Employers on this board attest they are hiring domestically.