AI Engineer, Evaluation
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.
Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.
What We Are Looking For
At Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.
AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.
This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.
Key Responsibilities
Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself
Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone
Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve
Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production
What We Require
5+ years of software engineering experience
Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code
Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes
Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production
AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office..
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
AI Engineer
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.
Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.
What We Are Looking For
AI Engineers build and operate production AI systems that deliver business value inside customer environments. This role is for engineers who thrive in ambiguous problem spaces, take ownership of outcomes, and want to work directly on AI systems that must perform reliably under enterprise constraints.
AI Engineers are hands-on builders. They design, implement, deploy, and iterate on end-to-end AI systems in close partnership with customers, subject matter experts, and other Distyl engineers. They translate messy operational needs into concrete system behavior, build the software and AI workflows required to support that behavior, and continuously improve systems through evaluation, feedback, integration, and production iteration.
This is not a demo-building role. AI Engineers are expected to make AI systems work in practice: with users, data, constraints, and accountability for production outcomes.
Key Responsibilities
Build and operate AI systems deployed in customer environments, taking ownership of system behavior, reliability, and usefulness in production
Design and implement compound AI workflows that combine models, prompts, agents, tools, retrieval, evaluation, feedback loops, and execution into coherent production systems aligned with user and SME needs
Develop clean, maintainable Python services and application logic that integrate AI capabilities into customer workflows, data platforms, APIs, and existing applications
Operate on live systems by measuring behavior, identifying failure modes, debugging issues, and iterating rapidly to improve quality, reliability, and user value
Build evaluation frameworks, test cases, feedback mechanisms, and observability patterns that help teams understand and improve AI system performance over time
Work directly with customer stakeholders and subject matter experts to understand workflows, clarify requirements, reason about tradeoffs, and adapt systems as needs evolve
Use AI-native engineering tools to accelerate implementation, debugging, experimentation, data analysis, and system improvement
Collaborate with other AI Engineers, AI Strategists, and other Distillers to make pragmatic system design decisions that balance speed, robustness, maintainability, and customer impact
Take accountability for the production outcomes of the components, workflows, and systems you build
What We Require
5+ years of software engineering experience
Ownership mentality for AI systems. You take responsibility for whether the systems you build deliver their intended value in production. You are comfortable making technical decisions, learning from system behavior, and owning the results of your work
Experience building AI systems. You have built applications powered by LLMs or other AI models and are comfortable composing multiple components — prompts, agents, tools, retrieval, evaluators, workflows, and integrations — into end-to-end systems. You reason about system behavior holistically rather than treating models as black boxes
Strong engineering fundamentals. You write clean, maintainable Python and are comfortable building production software systems. You understand core engineering concepts like versioning, debugging, testing, performance, code review, and production readiness
AI-native working style. You use AI tools daily to write and debug code, explore designs, analyze data, and automate repetitive work. You are curious about new model capabilities and techniques, and actively incorporate them into how you build and iterate on systems
Comfort in customer environments. You are able to work directly with customer teams, ask good questions, and adapt quickly to new domains. You communicate clearly about system behavior, limitations, and tradeoffs, and can operate effectively in high-trust, high-visibility situations
Pragmatic delivery mindset. You can navigate ambiguity, make progress with incomplete information, and balance speed with robustness when building systems that need to work for production users
Willingness to travel. Travel is typically 10–30%, depending on the project, customer needs, and your role on the engagement
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office.
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
AI Engineer, Enterprise Lead
AI Engineer, Lead
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.
Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.
What We Are Looking For
AI Engineers are hands-on technical leaders who own the architecture, execution, and delivery of critical production AI systems for large enterprise customers. They make technical decisions, lead project delivery, coach engineers, and partner closely with customer stakeholders and Distyl’s AI strategy team to ensure systems deliver real business value.
This role is for engineers who want ownership over production outcomes: shaping system architecture, leading technical execution, working directly with customers, and improving AI systems under real-world constraints.
Key Responsibilities
Lead and mentor a team of AI Engineers—owning technical direction, code quality, execution standards, and individual growth
Design and implement AI systems that combine models, agents, retrieval, evaluation, and execution into coherent, production-ready systems aligned with real business outcomes
Work directly with customer stakeholders, often in high-visibility settings, and communicate clearly about system behavior, tradeoffs, limitations, and paths to improvement
Operate and improve live AI systems by measuring behavior, identifying failure modes, debugging issues, and rapidly iterating on quality, reliability, and usefulness
Integrate AI systems into customer data platforms, APIs, and existing applications. Make pragmatic system design decisions that balance speed, robustness, maintainability, and long-term operability
Take accountability for outcomes in production and adapt systems as requirements evolve
Who You Are
8+ years of engineering experience, including experience as a tech lead or engineering lead on customer-facing or production AI projects
Ownership mentality for AI systems. You take responsibility for whether an AI system delivers its intended value in production. You are comfortable making independent technical decisions across system design, evaluation, integration, and iteration
Technical leadership in teams. Management experience is not required, but you should have led engineers through technical decision-making, execution, mentorship, and delivery. Growth in this role includes taking on broader technical and leadership scope over time
Strong solutions architecture fundamentals: You have experience with cloud systems, system integrations, API design, and data engineering. You can understand how an AI system fits into a broader enterprise ecosystem and operate as a peer to customer architecture and engineering teams
AI-Native Working Style: You use AI tools daily to write and debug code, explore designs, analyze data, and automate repetitive work. You are curious about new model capabilities and techniques, and actively incorporate them into how you build and iterate on systems
Willingness to travel: Travel is typically 10–30%, depending on the project, customer needs, and your role on the engagement
What We Offer
The base salary range for this role is $200-$250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office.
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
Software Engineer - Back End
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.
Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.
What We Are Looking For
As a Software Engineer Back End you will help design, build, and optimize Distillery—our AI-native platform that powers real-world enterprise AI systems for diverse F500 workflows.
Your role will involve developing scalable AI infrastructure, ensuring system reliability, and collaborating with engineers and business leaders to solve some of the most complex AI deployment challenges.
Key Responsibilities
Build & Scale AI-Native Infrastructure: Develop and refine a platform where AI builds, optimizes, and operates AI-powered workflows. Define how AI automation integrates into traditional enterprise infrastructure.
Develop Cloud-Native Microservices & Scalable AI Systems: Design and build secure, high-performance backend services deployed across AWS/GCP/Azure or on-prem Kubernetes environments. Build using Python, FastAPI, SQLAlchemy, Alembic, and modern DevOps tools to develop scalable, reliable AI infrastructure.
Optimize System Performance & Reliability: Ensure high-availability, security, and observability of AI-native workflows. Implement best practices for ML/AI Ops, distributed computing, and scalable service orchestration.
Collaborate with Cross-Functional Teams: Partner with Forward-Deployed Engineers (FDEs), AI Researchers, and business SMEs to translate real-world operational needs into platform capabilities. Advocate for strong software engineering and DevOps practices, driving high coding standards and scalable architectures.
Who You Are
We are hiring multiple roles across different levels of seniority (3-10+ years of software engineering experience)
Proficiency in Backend & Systems Engineering. Expertise in Python, Java, Golang, or C++ for building scalable, high-performance systems
Hands-on experience with Kubernetes, CI/CD, cloud platforms (AWS, GCP, or Azure), and infrastructure as code
Strong interest in AI-native development, leveraging tools like ChatGPT, Claude, Perplexity, and Cursor in engineering workflows
Experience with security, distributed systems, storage, ML/AI Ops, and large-scale observability is a plus
Travel: Ability to travel 10-20%
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office.
We’re grateful for the strong interest in this role. The best way to get your profile in front of our team is to apply directly through our careers page, where all applications are reviewed. Due to the high volume of interest, we’re not able to review or respond to all direct emails or LinkedIn messages. We will be in touch with every applicant once we’ve completed our review, regardless of the decision.
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
The full posting opens here — pay, setting and the full description, without leaving the list.