Research Engineer, Code Agents Infra
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
This role focuses on building and operating the end-to-end execution, training, and data infrastructure that powers Mistral’s agentic models and coding assistants. You will be a core contributor to our agent research stack: designing scalable systems for synthetic data generation, building ultra-fast training and RL execution environments, and maintaining high-throughput execution engines.
You will tackle the engineering challenges at every step of the agent lifecycle: from orchestrating 1M+ concurrent and short-lived sandboxes for untrusted code execution to optimizing agent training codebases, distributed trajectory collection pipelines, and dataset processing workflows across massive hybrid and multi-cloud clusters.
What You Will Do
Large-Scale Sandboxing Infrastructure: Design, deploy, and operate our high-throughput sandboxing platform, executing LLM-generated untrusted code across over 1 million isolated environments concurrently for model evaluation and interactive RL environments.
Agent Data Generation Pipelines: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops.
Training Codebase & Systems Optimization: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks.
Low-Latency Orchestration & Warm Pooling: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (e.g., CRIU, microVMs), and optimized image delivery layers across hybrid clusters.
Multi-Cluster Queueing & Resource Allocation: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.
Isolation, Security & Security Boundary: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (e.g., gVisor, Firecracker) and default-deny network postures.
Operational Excellence: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations for critical agent training and execution pipelines.
What We're Looking For
4+ years of experience in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads.
Data & Pipeline Engineering: Proven experience building high-throughput data processing and generation pipelines for large-scale datasets (e.g., Ray, Spark, custom distributed queues).
Deep experience with Kubernetes & Container Tech: Strong expertise writing custom K8s operators/controllers, managing Linux cgroups/namespaces, and optimizing Docker image layers and distribution systems.
High-Performance Software Engineering: Advanced proficiency in Python, Go, C++ or Rust, with a track record of profiling and optimizing high-performance ML or backend systems codebases.
Sandboxing & Isolation Technologies: Hands-on experience with lightweight virtualization, container runtimes, or WASM (e.g., Docker, gVisor, Firecracker).
Queueing & Scheduling: Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.
Comfort with Ambiguity: Passion for working directly alongside AI researchers to rapidly turn frontier agent ideas into scalable, production-grade infrastructure.
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our .
Research Engineer, Code Agents Infra
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
This role focuses on building and operating the end-to-end execution, training, and data infrastructure that powers Mistral’s agentic models and coding assistants. You will be a core contributor to our agent research stack: designing scalable systems for synthetic data generation, building ultra-fast training and RL execution environments, and maintaining high-throughput execution engines.
You will tackle the engineering challenges at every step of the agent lifecycle: from orchestrating 1M+ concurrent and short-lived sandboxes for untrusted code execution to optimizing agent training codebases, distributed trajectory collection pipelines, and dataset processing workflows across massive hybrid and multi-cloud clusters.
What You Will Do
Large-Scale Sandboxing Infrastructure: Design, deploy, and operate our high-throughput sandboxing platform, executing LLM-generated untrusted code across over 1 million isolated environments concurrently for model evaluation and interactive RL environments.
Agent Data Generation Pipelines: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops.
Training Codebase & Systems Optimization: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks.
Low-Latency Orchestration & Warm Pooling: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (e.g., CRIU, microVMs), and optimized image delivery layers across hybrid clusters.
Multi-Cluster Queueing & Resource Allocation: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.
Isolation, Security & Security Boundary: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (e.g., gVisor, Firecracker) and default-deny network postures.
Operational Excellence: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations for critical agent training and execution pipelines.
What We're Looking For
4+ years of experience in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads.
Data & Pipeline Engineering: Proven experience building high-throughput data processing and generation pipelines for large-scale datasets (e.g., Ray, Spark, custom distributed queues).
Deep experience with Kubernetes & Container Tech: Strong expertise writing custom K8s operators/controllers, managing Linux cgroups/namespaces, and optimizing Docker image layers and distribution systems.
High-Performance Software Engineering: Advanced proficiency in Python, Go, C++ or Rust, with a track record of profiling and optimizing high-performance ML or backend systems codebases.
Sandboxing & Isolation Technologies: Hands-on experience with lightweight virtualization, container runtimes, or WASM (e.g., Docker, gVisor, Firecracker).
Queueing & Scheduling: Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.
Comfort with Ambiguity: Passion for working directly alongside AI researchers to rapidly turn frontier agent ideas into scalable, production-grade infrastructure.
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our .
Applied Scientist
Applied Scientist
Research Engineer, Data Infrastructure
Site Reliability Engineer
Software Engineer, Backend
Research Engineer, Machine Learning
Applied AI, Technical Lead - Forward Deployed AI Engineer
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
Mistral AI is seeking a Technical Lead, Applied AI to drive the technical strategy, execution, and delivery of complex AI solutions for our enterprise customers. In this role, you will lead a project teams of Applied AI Engineers, ensuring the successful deployment of Mistral AI products and the development of high-impact, scalable AI use cases.
You will act as the primary technical point of contact for our most strategic customers, guiding them through the entire lifecycle—from pre-sales to post-implementation—while collaborating closely with research, product, and engineering teams to shape the future of our offerings.
As a Technical Lead, you will bridge the gap between cutting-edge AI research and real-world enterprise applications, ensuring our solutions are robust, scalable, and aligned with both customer needs and Mistral’s technological vision.
What You Will Do
- Deliver as an IC the critical lines of codes of our complex projects, you’ll be hands-on and de-risk the critical parts of our complex projects. You’ll stay deeply involved in coding, reviewing, and optimizing AI solutions.
- Lead technical teams of Applied AI Engineers, providing mentorship, technical guidance, and best practices for deploying state-of-the-art GenAI applications across industries.
- Lead technical discussions during pre-sales, translating customer requirements into actionable solutions and communicating Mistral’s technological advantages to diverse stakeholders.
- Design and oversee the implementation of complex AI systems, including fine-tuning, RAG, agentic workflows, and custom LLM applications, ensuring alignment with Mistral’s product roadmap and open-source initiatives.
- Drive innovation by identifying emerging trends in AI, evaluating new tools and methodologies, and championing best practices for fine-tuning, inference, and deployment.
- Work closely with product managers, researchers, and engineers to ensure seamless integration of customer feedback into Mistral’s product development cycle.
What We're Looking For
- You hold a PhD or Master’s degree in AI, Machine Learning, Computer Science, or a related field.
- You have 7/8+ years of experience in AI/ML, with at least 2+ years in a technical leadership role (e.g., Tech Lead, Engineering Manager, or Solutions Architect) focused on AI products or enterprise solutions.
- You have a proven track record of leading teams to deliver complex AI projects, from prototyping to production, in industries such as tech, finance, healthcare, or industrial automation.
- You possess deep expertise in fine-tuning LLMs, advanced RAG, agentic systems, and deploying NLP applications at scale.
- You are proficient in Python, PyTorch, and modern AI frameworks (e.g., LangChain, Hugging Face). Experience with cloud platforms (AWS, GCP, Azure) and MLOps tools is a plus.
- You have strong software engineering skills, including API design, backend/full-stack development, and system architecture.
- You excel in technical communication, with the ability to articulate complex concepts to both technical and non-technical audiences, including executives and engineers.
- You thrive in fast-paced, collaborative environments and are passionate about mentoring and growing technical talent.
Ideally, you have:
- Contributed to open-source projects, particularly in the LLM or AI space.
- Experience in customer-facing roles (e.g., Solutions Architect, Customer Engineer, or Technical Product Manager) with a focus on enterprise AI adoption.
- A track record of driving technical strategy and influencing product direction based on customer needs and market opportunities.
Why join us? You’ll have the opportunity to shape the future of AI adoption in enterprises, work with a world-class team, and contribute to open-source projects that impact millions. If you’re excited about leading technical innovation and solving real-world challenges with AI, we’d love to hear from you!
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our .
Site Reliability Engineer
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and exceed our internal and external customers' expectations.
What You Will Do
As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.
• Design, build, and maintain scalable, highly available and fault-tolerant infrastructures to support our web services and ML workloads
• Make sure our platform, inference and model training environments are always highly available and enable seamless replication of work environments across several HPC clusters
• Operate systems and troubleshoot issues in production environments (interrupts, on-call responses, users admin, data extraction, infrastructure scaling, etc.)
• Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime
• Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our client-facing APIs and large training runs
• Participate occasionally in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences
• Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
• Collaborate with AI/ML researchers to develop and implement solutions that enable safe and reproducible model-training experiments
• Build a cloud-agnostic platform offering an abstraction layer between science and infrastructure
• Design and develop new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.)
• Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements
• Document processes and procedures to ensure consistency and knowledge sharing across the team
• Contribute to open-source projects, research publications, blog articles and conferences
What We're Looking For
• Master’s degree in Computer Science, Engineering or a related field
• 7+ years of experience in a DevOps/SRE role
• Strong experience with cloud computing and highly available distributed systems
• Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...)
• Experience working against reliability KPIs (observability, alerting, SLAs)
• Hands-on experience with CI/CD, containerization and orchestration tools (Docker, Kubernetes...)
• Knowledge of monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack, Datadog...)
• Familiarity with infrastructure-as-code tools like Terraform or CloudFormation
• Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices
• Strong understanding of networking, security, and system administration concepts
• Excellent problem-solving and communication skills
• Self-motivated and able to work well in a fast-paced startup environment
Your application will be all the more interesting if you also have:
• experience in an AI/ML environment
• experience of high-performance computing (HPC) systems and workload managers (Slurm)
• worked with modern AI-oriented solutions (Fluidstack, Coreweave, Vast...)
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our .
The full posting opens here — pay, setting and the full description, without leaving the list.