Page 1
Research Scientist - Humanoid Robotics
Applied Intuition · Sunnyvale, California, United States
$250k–423k
18 days ago
♡
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
$140k–210k
21 days ago
♡
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
$180k–500k
22 days ago
♡
Research Scientist I/II, Immunology
Iambic Therapeutics · San Diego, California, United States
$113k–156k
23 days ago
♡
Loading more openings…
You've reached the end of the list.
Research Scientist - Humanoid Robotics
Applied Intuition · Sunnyvale, California, United States
Pay
$250k–423k
Setting
On-site
Listed by Applied Intuition for a position based in the United States. Employers on this board attest they are hiring domestically.
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
Pay
$140k–210k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Product Operations Manager, Research Services, you'll drive operational excellence to support Mercor's full-service research partnerships. You will prioritize and manage delivery across various strategic customer segments to support model evaluations, diagnose model capability gaps, and fulfill post-training demand to better support our most strategic customers. This is a highly technical, internally-facing role in the Product Operations job family at the intersection of Product, Engineering, Data Production, GTM, Finance, and Strategic Partnerships.
WHAT YOU'LL DO
- Prioritize the Benchmark Portfolio: Own the intake and prioritization of new benchmarks and evaluations supported in our platform. Weigh customer demand, observed capability gaps, differentiation versus public benchmarks, and build cost to inform our priorities.
- Support Dataset Demand Generation: Run and share leaderboard evals to drive Mercor’s research brand equity during new model releases.
- Manage Training and Inference Partners & Budgets: Own the relationships with our inference providers, compute partners, and external training vendors. Forecast capacity against the delivery calendar, negotiate and track SLAs and rate cards, monitor cost per eval run and per rollout, and qualify new partners before we need them.
- Track Customer Delivery: Maintain the operating picture across every active evals, loss analysis, and post-training engagements, highlighting scope, milestones, dependencies, and SLA status. Run the delivery review, escalate slipping commitments before the customer notices, and give research counterparts a straight answer on where things stand.
- Shape the Product: Translate recurring delivery friction into product requirements for the eval platform and partner with Engineering to automate the workflows you'd otherwise be running by hand.
WHAT WE'RE LOOKING FOR
- Experience: 3+ years in Product Operations, Product Management, Technical Program Management, Research Operations, Solutions Engineering, or a similar technical, customer-facing role. Experience operating in an ML research or AI infrastructure environment strongly preferred.
- Technical Fluency: Fluent in how modern LLM evaluation works, including harness design, verification, agentic rollouts, and contamination and fairness pitfalls. Working understanding of post-training (SFT, preference ranking, RL) and of inference economics. Familiar with coding agents; comfortable in SQL and Python without agent assistance.
- Vendor and Capacity Judgment: Can plan capacity against an uncertain demand curve, hold partners to an SLA, and weigh cost against reliability without needing to escalate every call.
- Ownership: Thrive in ambiguous environments and take ownership from problem definition through execution. Comfortable being the person accountable for a date.
- Communication: Can hold your own in a technical conversation with AI researchers and write a status update an executive can act on. Excellent written and verbal communication.
- Systems Thinking: Strong project management and cross-functional coordination skills. Passion for fixing problems at the source and building repeatable systems rather than one-off solutions.
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Research Data Engineer
Canva · Vienna, Vienna, AT
Pay
Not listed
Setting
On-site
Listed by Canva for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
Pay
$180k–500k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Research Engineer at Mercor, you’ll work at the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how we train and improve frontier language models.
Your work will define how we measure tool use, agentic behavior, and real-world reasoning. You’ll design and run evals, build rubrics and scorers, and turn failure analysis into actionable improvements for post-training, RLVR, and data pipelines.
WHAT YOU’LL DO
- Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals.
- Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale.
- Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
- Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement).
- Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation.
- Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most.
- Ownership in a fast-paced environment: Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows.
WHAT WE’RE LOOKING FOR
- Strong applied research background, with focus on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid grasp of data structures, algorithms, and backend systems.
- Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
- Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses.
- Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership environment.
NICE TO HAVE
- Industry experience on a post-training or evaluation/benchmarking team (highest priority).
- Publications at top-tier venues (NeurIPS, ICML, ACL), especially in evaluation or benchmarking.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
- Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping.
- Work samples or code (e.g., eval frameworks, benchmark suites, failure-analysis reports or tooling) that demonstrate relevant skills.
BENEFITS
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Data Scientist
Guidehouse · VA, McLean, US
Pay
$98k–163k
Setting
On-site
Listed by Guidehouse for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Scientist I/II, Immunology
Iambic Therapeutics · San Diego, California, United States
Pay
$113k–156k
Setting
On-site
Listed by Iambic Therapeutics for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer, Lab Automation
Periodic Labs · Menlo Park, California, United States
Pay
$200k–250k
Setting
On-site
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.
About the Role
Join our team of scientists and engineers building a lab where AI and automation speed up materials discovery.
We're building an autonomous lab to speed up materials discovery, and we need someone who understands materials R&D from the inside — someone who's run the experiments, fought with the instruments, and knows what "the workflow" actually means at the bench. As our Research Engineer, you'll work directly with scientists to understand what they're trying to learn, then translate that into the hardware setups, instrument sequences, and engineering requirements that make it possible to automate.
What You'll Do
- Work closely with bench scientists to understand experimental goals and turn them into concrete hardware and workflow requirements.
- Evaluate, select, and configure lab instrumentation and hardware to support new and existing materials R&D workflows.
- Design experimental and automation workflows that hold up to the realities of materials synthesis and characterization.
- Serve as the technical bridge between scientists and the automation/software engineering team, making sure integration specs reflect how the science actually works.
- Troubleshoot instrument and workflow issues that require materials domain knowledge to diagnose.
- Collaborate with AI and data scientists to help shape how experimental data is structured and used for analysis and planning.
You Will Thrive in This Role If You Have
- PhD in Materials Science, Chemistry, Chemical Engineering, or a related field (or equivalent research experience).
- Strong working knowledge of common materials lab hardware (e.g. furnaces, fluid/gas handling manifolds, characterization tools, synthesis equipment) and the workflows built around them.
- Ability to communicate fluently with both scientists and engineers, and to translate scientific intent into clear technical requirements.
- Comfort writing Python to interact with instruments, manipulate experimental data, and prototype automation workflows
Especially Strong Candidates May Also Have
- Prior experience working alongside automation or software engineers.
- Experience with electronic lab notebooks, LIMS, or other lab data systems.
- A track record of designing or adapting experimental protocols for higher-throughput or automated execution.
Mechanics
- Minimum education: PhD or equivalent combination of education and hands-on research experience
- Location: Menlo Park, CA (Soon: San Francisco, too)
- Compensation: $200,000-$250,000 + equity
- Visa sponsorship: Yes, we sponsor visas.
Listed by Periodic Labs for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Staff, Clinical Microbiology
Guidehouse · MD, Bethesda, US
Pay
$149k–248k
Setting
On-site
Listed by Guidehouse for a position based in the United States. Employers on this board attest they are hiring domestically.
Associate Public Health Informaticist
Guidehouse · TX, San Antonio, US
Pay
$53k–88k
Setting
On-site
Listed by Guidehouse for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer - Midtraining
Periodic Labs · Menlo Park, California, United States
Pay
$250k–350k
Setting
On-site
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.
ABOUT THE ROLE
We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.
WHAT YOU'LL DO
- Identify, process, and curate novel sources of scientific data for large-scale model training.
- Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.
- Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.
- Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.
- Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.
- Build tools for yourself and the team to investigate how data choices shape model intelligence.
YOU WILL THRIVE IN THIS ROLE IF YOU HAVE
- Experience training LLMs on curated mixes of trillions of tokens.
- Experience on a dedicated evals team supporting a large production training run.
- Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.
- Experience with scaling laws and compute-optimal hyperparameters.
- Comfort working across data, evals, and training infrastructure.
ESPECIALLY STRONG CANDIDATES MAY ALSO HAVE
- Experience optimizing throughput and reliability for large-scale distributed training runs.
- A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).
- Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.
MECHANICS
- Minimum education: Bachelor's degree or similar experience
- Location: Menlo Park, CA (Soon: San Francisco, too)
- Compensation: $250,000–$350,000 + equity
- Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.
Listed by Periodic Labs for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.