Page 1
Mercor Research Fellowship — APEX
Mercor · San Francisco, California, United States
$40k
18 days ago
♡
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
$140k–210k
21 days ago
♡
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
$180k–500k
22 days ago
♡
Research Engineer, Real Environments
Mercor · San Francisco, California, United States
$180k–500k
47 days ago
♡
Loading more openings…
You've reached the end of the list.
Mercor Research Fellowship — APEX
Mercor · San Francisco, California, United States
Pay
$40k
Setting
Remote
Location: Remote. Fellows can be hosted on-site in San Francisco. Employment Type: Fellowship, fixed-term (3–6 months) Department: Research (APEX)
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
Pay
$140k–210k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Product Operations Manager, Research Services, you'll drive operational excellence to support Mercor's full-service research partnerships. You will prioritize and manage delivery across various strategic customer segments to support model evaluations, diagnose model capability gaps, and fulfill post-training demand to better support our most strategic customers. This is a highly technical, internally-facing role in the Product Operations job family at the intersection of Product, Engineering, Data Production, GTM, Finance, and Strategic Partnerships.
WHAT YOU'LL DO
- Prioritize the Benchmark Portfolio: Own the intake and prioritization of new benchmarks and evaluations supported in our platform. Weigh customer demand, observed capability gaps, differentiation versus public benchmarks, and build cost to inform our priorities.
- Support Dataset Demand Generation: Run and share leaderboard evals to drive Mercor’s research brand equity during new model releases.
- Manage Training and Inference Partners & Budgets: Own the relationships with our inference providers, compute partners, and external training vendors. Forecast capacity against the delivery calendar, negotiate and track SLAs and rate cards, monitor cost per eval run and per rollout, and qualify new partners before we need them.
- Track Customer Delivery: Maintain the operating picture across every active evals, loss analysis, and post-training engagements, highlighting scope, milestones, dependencies, and SLA status. Run the delivery review, escalate slipping commitments before the customer notices, and give research counterparts a straight answer on where things stand.
- Shape the Product: Translate recurring delivery friction into product requirements for the eval platform and partner with Engineering to automate the workflows you'd otherwise be running by hand.
WHAT WE'RE LOOKING FOR
- Experience: 3+ years in Product Operations, Product Management, Technical Program Management, Research Operations, Solutions Engineering, or a similar technical, customer-facing role. Experience operating in an ML research or AI infrastructure environment strongly preferred.
- Technical Fluency: Fluent in how modern LLM evaluation works, including harness design, verification, agentic rollouts, and contamination and fairness pitfalls. Working understanding of post-training (SFT, preference ranking, RL) and of inference economics. Familiar with coding agents; comfortable in SQL and Python without agent assistance.
- Vendor and Capacity Judgment: Can plan capacity against an uncertain demand curve, hold partners to an SLA, and weigh cost against reliability without needing to escalate every call.
- Ownership: Thrive in ambiguous environments and take ownership from problem definition through execution. Comfortable being the person accountable for a date.
- Communication: Can hold your own in a technical conversation with AI researchers and write a status update an executive can act on. Excellent written and verbal communication.
- Systems Thinking: Strong project management and cross-functional coordination skills. Passion for fixing problems at the source and building repeatable systems rather than one-off solutions.
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
Pay
$180k–500k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Research Engineer at Mercor, you’ll work at the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how we train and improve frontier language models.
Your work will define how we measure tool use, agentic behavior, and real-world reasoning. You’ll design and run evals, build rubrics and scorers, and turn failure analysis into actionable improvements for post-training, RLVR, and data pipelines.
WHAT YOU’LL DO
- Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals.
- Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale.
- Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
- Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement).
- Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation.
- Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most.
- Ownership in a fast-paced environment: Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows.
WHAT WE’RE LOOKING FOR
- Strong applied research background, with focus on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid grasp of data structures, algorithms, and backend systems.
- Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
- Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses.
- Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership environment.
NICE TO HAVE
- Industry experience on a post-training or evaluation/benchmarking team (highest priority).
- Publications at top-tier venues (NeurIPS, ICML, ACL), especially in evaluation or benchmarking.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
- Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping.
- Work samples or code (e.g., eval frameworks, benchmark suites, failure-analysis reports or tooling) that demonstrate relevant skills.
BENEFITS
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer, Real Environments
Mercor · San Francisco, California, United States
Pay
$180k–500k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
You’ll work with large enterprises to capture their data and transform it into high-fidelity RL environments for capability evaluations and training datasets for frontier labs. We focus on pushing the frontier of world-building, verifier engineering, and more alongside our partners.
Your goal will be to automate the process of building evals for real work in the economy.
WHAT YOU'LL DO
- Ship models for workflow extraction, classification, and grading.
- Engineer autonomous task refinement processes which distill data taste into pipelines.
- Deliver data to customers and deploy into real engagements.
- Help define the future of agentic transformation for enterprises around the world.
- Deeply learn about the intricacies of enterprises through building evaluations for all aspects of work.
- Build end-to-end environments for labs & enterprises by platformizing sandbox app clones, load real data into the sandboxes, build prompts from real workflows, and write verifiers leveraging enterprise expertise & golden outputs.
- Systematize the production of environments to scale throughput while maintaining high-quality worlds and verifiers.
WHAT WE'RE LOOKING FOR
- Prior experience shipping environments – you’ve contributed to an OSS framework, built environments at previous companies, or worked on agentic evaluations.
- Strong full-stack engineering skills – you’ll be responsible for everything from infrastructure to app code to analytics
- Bias to action – this team is focused on shipping evals, not just philosophizing about them.
- Curiosity – being biased towards understanding and digging deep into model behavior and actually looking at the data.
- Sweat the details that make a simulation indistinguishable from the real thing and have systems-level thinking skills that allow you to scale up quality.
NICE TO HAVE
- Experience with Temporal, Modal, or similar orchestration/compute services
- Experience with synthetic data generation for frontier models.Past work auditing and scrutinizing industry-standard evaluations
BENEFITS
- Semi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Operations, Code
Mercor · San Francisco, California, United States
Pay
$130k–250k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
In the Research Operations role at Mercor, you will work at the intersection of engineering and applied AI research. This role will own multi-million-dollar data programs end to end with responsibilities to:
- Reason about opportunities for frontier coding models, and translate that into the data that will make it stronger.
- Shape how data gets produced, including human-expert workflows, synthetic generation pipelines, and model-in-the-loop systems.
- Turn designs into scaled, high-quality data pipelines and own deliveries to customers.
- Build deep relationships with lab researchers and become the person they trust to tell them what data they actually need.
WHAT WE'RE LOOKING FOR
- Bachelor’s degree in Computer Science, a STEM field, or an equivalent technical discipline.
- At least two years of experience in software development or a forward deployed engineering role.
- A track record of running complex, high-stakes projects end to end, and genuine energy for large-scale execution and gritty process optimization under pressure.
- Strong analytical and communication skills; comfortable owning high-profile relationships with technical customers.
NICE TO HAVE
- Experience designing scalable data or automation pipelines
- Research fluency: connecting model/benchmark literature to what data would move a frontier model; having created a benchmark or published analysis of model behavior.
- Backgrounds that often fit: data scientists, software engineers, or forward deployed engineers who love operating, technical PMs, research engineers, or strong generalist operators with real technical range. Backgrounds from consulting, finance, or high-growth startups can work if paired with real technical fluency.
BENEFITS
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K proximity bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer - Environments, Data and Post-Training
Mercor · San Francisco, California, United States
Pay
$180k–500k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Research Engineer at Mercor, you’ll work at the intersection of engineering and applied AI research. You’ll contribute directly to post-training and RLVR, synthetic data generation, and large-scale evaluation workflows that meaningfully impact frontier language models.
Your work will be used to train large language models to master tool use, agentic behavior, and real-world reasoning in real-world production environments. You’ll shape rewards, run post-training experiments, and build scalable systems that improve model performance. You’ll help design and evaluate datasets, create scalable data augmentation pipelines, and build rubrics and evaluators that push the boundaries of what LLMs can learn.
WHAT YOU’LL DO
- Work on post-training and RLVR pipelines to understand how datasets, rewards, and training strategies impact model performance.
- Design and run reward-shaping experiments and algorithmic improvements (e.g., GRPO, DAPO) to improve LLM tool-use, agentic behavior, and real-world reasoning.
- Quantify data usability, quality, and performance uplift on key benchmarks.
- Build and maintain data generation and augmentation pipelines that scale with training needs.
- Create and refine rubrics, evaluators, and scoring frameworks that guide training and evaluation decisions.
- Build and operate LLM evaluation systems, benchmarks, and metrics at scale.
- Collaborate closely with AI researchers, applied AI teams, and experts producing training data.
- Operate in a fast-paced, experimental research environment with rapid iteration cycles and high ownership.
WHAT WE’RE LOOKING FOR
- Strong applied research background, with a focus on post-training and/or model evaluation.
- Strong coding proficiency and hands-on experience working with machine learning models.
- Strong understanding of data structures, algorithms, backend systems, and core engineering fundamentals.
- Familiarity with APIs, SQL/NoSQL databases, and cloud platforms.
- Ability to reason deeply about model behavior, experimental results, and data quality.
- Excitement to work in person in San Francisco, five days a week (with optional remote Saturdays), and thrive in a high-intensity, high-ownership environment.
NICE TO HAVE
- Real-world post-training team experience in industry (highest priority).
- Publications at top-tier conferences (NeurIPS, ICML, ACL).
- Experience training models or evaluating model performance.
- Experience in synthetic data generation, LLM evaluations, or RL-style workflows.
- Work samples, artifacts, or code repositories demonstrating relevant skills.
BENEFITS
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.