Page 1
Senior ML Research Scientist
Rad AI · San Francisco, California, United States
$170k–220k
12 days ago
♡
Research Scientist - Humanoid Robotics
Applied Intuition · Sunnyvale, California, United States
$250k–423k
17 days ago
♡
Mercor Research Fellowship — APEX
Mercor · San Francisco, California, United States
$40k
17 days ago
♡
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
$140k–210k
20 days ago
♡
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
$180k–500k
21 days ago
♡
Loading more openings…
You've reached the end of the list.
Senior ML Research Scientist
Rad AI · San Francisco, California, United States
Pay
$170k–220k
Setting
Remote
ABOUT RAD AI
At Rad AI, we’re on a mission to transform healthcare with artificial intelligence. Founded by a radiologist, our AI-driven solutions are revolutionizing radiology—saving time, reducing burnout, and improving patient care. With one of the largest proprietary radiology report datasets in the world, our AI has helped uncover hundreds of new cancer diagnoses and reduced error rates in tens of millions of radiology reports by nearly 50%.
Rad AI has secured over $140M in funding, including a recently oversubscribed Series C ($68M round) led by Transformation Capital, bringing our valuation to $528M. Our investors include Khosla Ventures, World Innovation Lab, Gradient Ventures, Cone Health Ventures, and others—all backing our mission to empower physicians with cutting-edge AI.
Our latest advancements in generative AI are used by thousands of radiologists daily, supporting more than one-third of radiology groups and healthcare systems and nearly 50% of all medical imaging in the U.S. at partners including Cone Health, Jefferson Einstein Health, Geisinger, Guthrie Healthcare System, and Henry Ford Health.
Recognized as one of the most promising healthcare AI companies by CB Insights and AuntMinnie https://www.radai.com/news/auntminnie-recognizes-rad-ai-omni-reporting-as-2023s-best-new-radiology-software, and ranked by Deloitte https://www2.deloitte.com/us/en/pages/technology-media-and-telecommunications/articles/fast500-winners.html as the 19th fastest-growing company in North America, we are building AI-powered solutions that make a real impact. Most recently, Rad AI was named to CNBC’s Disruptor 50 https://www.cnbc.com/2025/06/10/2025-cnbc-disruptor-50-see-the-full-list-of-companies.html list, highlighting the innovation and momentum behind our mission.
If you’re ready to shape the future of healthcare, we’d love to have you on our team!
WHAT YOU’LL DO
- Own a multimodal ML work-stream from problem definition through experimentation, evaluation, deployment, and iteration.
- Translate clinical and product needs into clear ML objectives, data strategies, model approaches, and success criteria.
- Build and evaluate modern ML systems, including transformers, self-supervised learning, weak supervision, detection, localization, and segmentation.
- Work with image, report, and other clinical data to develop systems that are useful in real radiology workflows.
- Design rigorous evaluations that go beyond aggregate offline metrics, including clinically meaningful operating points, robustness, calibration, and performance across relevant data slices.
- Partner with engineering to productionize models, make practical system tradeoffs, and learn from performance after launch.
- Investigate failure modes such as laterality errors, poor image or report grounding, hallucination, dataset bias, domain shift, and workflow disruption.
- Communicate research findings and technical decisions clearly through design documents, experiment reviews, and presentations to technical and clinical partners.
- Contribute to the research roadmap by identifying promising approaches, sharing learnings, and helping the team decide what to pursue next.
- Mentor less experienced researchers and engineers through project collaboration, code and experiment reviews, and technical guidance.
WHAT WE’RE LOOKING FOR
- Strong applied experience in computer vision, NLP, or deep learning, with a track record of independently designing experiments, analyzing results, and turning findings into working systems.
- Experience owning substantial ML projects across the full lifecycle, from data and modeling through production delivery.
- Deep hands-on ability in Python and PyTorch, with strong intuition for model architecture, data quality, experimentation, and evaluation.
- Experience with modern vision or multimodal techniques such as vision transformers, contrastive learning, masked image modeling, or weak supervision, etc.
- The judgment to connect model performance to real user and clinical outcomes, including knowing when a benchmark improvement is not enough.
- Strong collaboration skills across research, engineering, product, data, and clinical teams.
- Clear written and verbal communication, including the ability to explain technical tradeoffs to both ML experts and clinical partners.
- Typically 4+ years of relevant applied ML research or engineering experience, or equivalent scope and impact. We calibrate on demonstrated ownership rather than title or exact tenure.
- An MS, PhD, or equivalent practical experience in Computer Science, Electrical Engineering, Machine Learning, Biomedical Engineering, or a related quantitative field.
NICE TO HAVE
- Experience with medical imaging, radiology, healthcare, or another high-stakes application area.
- Familiarity with chest X-ray, CT, MRI, mammography, or other clinical imaging modalities.
- Experience with DICOM, image-report pairing, medical data de-identification, radiology workflows, or clinically derived labels.
- Experience evaluating models across patients, sites, scanner vendors, protocols, or other sources of distribution shift.
- Familiarity with clinical validation, FDA or HIPAA considerations, or other regulated and privacy-sensitive environments.
- Experience with 3D vision, longitudinal imaging, report generation, or clinical decision support.
- Publications, open-source contributions, or other evidence of research credibility.
WHAT SUCCESS LOOKS LIKE
You’ll own and advance a meaningful research track from ideation through production. You’ll establish a strong understanding of the clinical problem, build a credible data and evaluation strategy, deliver models that perform reliably in practice, and help the team learn from real-world use.
You’ll also become a trusted technical partner to the researchers, engineers, product leaders, data teams, and clinicians working on the broader ML roadmap. Over time, you’ll help raise the quality of research and technical decision-making through strong experimentation, clear communication, and thoughtful mentorship.
OUR WORKING STYLE
We’re a remote-first company with a highly collaborative, mission-driven research and engineering culture. We value direct communication, intellectual honesty, strong ownership, and practical judgment. The best work here comes from people who can go deep technically, stay close to the clinical context, and make progress even when the problem and the path are not fully defined.
This role is U.S. remote, with San Francisco Bay Area preferred. We encourage people from a wide range of backgrounds to apply. If the scope of this role excites you but your experience does not match every bullet, we would still love to hear from you.
Join our world-class team as we build and deploy AI solutions that empower physicians and transform patient care—making a meaningful impact on millions of lives. Driven by our mission, we prioritize transparency, inclusion, and close collaboration, bringing together exceptional people to revolutionize healthcare. If you're passionate about driving innovation and delivering impactful healthcare solutions, we'd love to hear from you!
To learn more about what it's like to work at Rad AI, visit https://www.radai.com/life-at-rad-ai and be sure to follow us on LinkedIn https://www.linkedin.com/company/radai/?utm_campaign=Recruiting_2026&utm_content=recruiting-job-posting-website&utm_source=recru%5B%E2%80%A6%5Db-posting-website to stay up to date!
Location Details:
For roles listed as San Francisco - Onsite:
- This role will be based in our San Francisco office and we expect employees to work onsite four days per week. The remaining time may be worked remotely or onsite, depending on team and business needs.
For roles listed as United States - Remote:
- This role is open to candidates located anywhere in the United States.
For roles listed as San Francisco - Onsite + United States - Remote:
- We will prioritize candidates who can work onsite four days per week in San Francisco, while also considering remote candidates located anywhere in the United States.
For US-Based Full-Time Roles, Rad AI offers a variety of benefits, including:
- Comprehensive Medical, Dental, Vision & Life insurance
- HSA (with employer match), FSA, & DCFSA
- 401(k)
- 11 Paid Company Holidays
- Flexible PTO policy
- Annual company-wide offsite
- Periodic team offsites
- Annual equipment stipend
- For roles based outside the US, your recruiter can share more details
At Rad AI, we value diversity and provide equal employment opportunities (EEO) to all employees and applicants without regard to race, color, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the San Francisco Fair Chance Ordinance.
Please be vigilant regarding job scams. We advise all candidates to apply directly through our official careers page. Our recruiters will use email addresses with the domain @radai.com http://radai.com or no-reply@ashbyhq.com.
Listed by Rad AI for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Scientist - Humanoid Robotics
Applied Intuition · Sunnyvale, California, United States
Pay
$250k–423k
Setting
On-site
Listed by Applied Intuition for a position based in the United States. Employers on this board attest they are hiring domestically.
Mercor Research Fellowship — APEX
Mercor · San Francisco, California, United States
Pay
$40k
Setting
Remote
Location: Remote. Fellows can be hosted on-site in San Francisco. Employment Type: Fellowship, fixed-term (3–6 months) Department: Research (APEX)
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Product Operations Manager, Research Services
Mercor · San Francisco, California, United States
Pay
$140k–210k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Product Operations Manager, Research Services, you'll drive operational excellence to support Mercor's full-service research partnerships. You will prioritize and manage delivery across various strategic customer segments to support model evaluations, diagnose model capability gaps, and fulfill post-training demand to better support our most strategic customers. This is a highly technical, internally-facing role in the Product Operations job family at the intersection of Product, Engineering, Data Production, GTM, Finance, and Strategic Partnerships.
WHAT YOU'LL DO
- Prioritize the Benchmark Portfolio: Own the intake and prioritization of new benchmarks and evaluations supported in our platform. Weigh customer demand, observed capability gaps, differentiation versus public benchmarks, and build cost to inform our priorities.
- Support Dataset Demand Generation: Run and share leaderboard evals to drive Mercor’s research brand equity during new model releases.
- Manage Training and Inference Partners & Budgets: Own the relationships with our inference providers, compute partners, and external training vendors. Forecast capacity against the delivery calendar, negotiate and track SLAs and rate cards, monitor cost per eval run and per rollout, and qualify new partners before we need them.
- Track Customer Delivery: Maintain the operating picture across every active evals, loss analysis, and post-training engagements, highlighting scope, milestones, dependencies, and SLA status. Run the delivery review, escalate slipping commitments before the customer notices, and give research counterparts a straight answer on where things stand.
- Shape the Product: Translate recurring delivery friction into product requirements for the eval platform and partner with Engineering to automate the workflows you'd otherwise be running by hand.
WHAT WE'RE LOOKING FOR
- Experience: 3+ years in Product Operations, Product Management, Technical Program Management, Research Operations, Solutions Engineering, or a similar technical, customer-facing role. Experience operating in an ML research or AI infrastructure environment strongly preferred.
- Technical Fluency: Fluent in how modern LLM evaluation works, including harness design, verification, agentic rollouts, and contamination and fairness pitfalls. Working understanding of post-training (SFT, preference ranking, RL) and of inference economics. Familiar with coding agents; comfortable in SQL and Python without agent assistance.
- Vendor and Capacity Judgment: Can plan capacity against an uncertain demand curve, hold partners to an SLA, and weigh cost against reliability without needing to escalate every call.
- Ownership: Thrive in ambiguous environments and take ownership from problem definition through execution. Comfortable being the person accountable for a date.
- Communication: Can hold your own in a technical conversation with AI researchers and write a status update an executive can act on. Excellent written and verbal communication.
- Systems Thinking: Strong project management and cross-functional coordination skills. Passion for fixing problems at the source and building repeatable systems rather than one-off solutions.
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Research Data Engineer
Canva · Vienna, Vienna, AT
Pay
Not listed
Setting
On-site
Listed by Canva for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer – Benchmarking
Mercor · San Francisco, California, United States
Pay
$180k–500k
Setting
On-site
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
As a Research Engineer at Mercor, you’ll work at the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how we train and improve frontier language models.
Your work will define how we measure tool use, agentic behavior, and real-world reasoning. You’ll design and run evals, build rubrics and scorers, and turn failure analysis into actionable improvements for post-training, RLVR, and data pipelines.
WHAT YOU’LL DO
- Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals.
- Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale.
- Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
- Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement).
- Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation.
- Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most.
- Ownership in a fast-paced environment: Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows.
WHAT WE’RE LOOKING FOR
- Strong applied research background, with focus on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid grasp of data structures, algorithms, and backend systems.
- Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
- Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses.
- Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership environment.
NICE TO HAVE
- Industry experience on a post-training or evaluation/benchmarking team (highest priority).
- Publications at top-tier venues (NeurIPS, ICML, ACL), especially in evaluation or benchmarking.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
- Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping.
- Work samples or code (e.g., eval frameworks, benchmark suites, failure-analysis reports or tooling) that demonstrate relevant skills.
BENEFITS
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Listed by Mercor for a position based in the United States. Employers on this board attest they are hiring domestically.
Data Scientist
Guidehouse · VA, McLean, US
Pay
$98k–163k
Setting
On-site
Listed by Guidehouse for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Scientist I/II, Immunology
Iambic Therapeutics · San Diego, California, United States
Pay
$113k–156k
Setting
On-site
Listed by Iambic Therapeutics for a position based in the United States. Employers on this board attest they are hiring domestically.
Lead Computational Social Scientist
Interos · Remote
Pay
Not listed
Setting
Remote
Listed by Interos for a position based in the United States. Employers on this board attest they are hiring domestically.
Research Engineer, Lab Automation
Periodic Labs · Menlo Park, California, United States
Pay
$200k–250k
Setting
On-site
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.
About the Role
Join our team of scientists and engineers building a lab where AI and automation speed up materials discovery.
We're building an autonomous lab to speed up materials discovery, and we need someone who understands materials R&D from the inside — someone who's run the experiments, fought with the instruments, and knows what "the workflow" actually means at the bench. As our Research Engineer, you'll work directly with scientists to understand what they're trying to learn, then translate that into the hardware setups, instrument sequences, and engineering requirements that make it possible to automate.
What You'll Do
- Work closely with bench scientists to understand experimental goals and turn them into concrete hardware and workflow requirements.
- Evaluate, select, and configure lab instrumentation and hardware to support new and existing materials R&D workflows.
- Design experimental and automation workflows that hold up to the realities of materials synthesis and characterization.
- Serve as the technical bridge between scientists and the automation/software engineering team, making sure integration specs reflect how the science actually works.
- Troubleshoot instrument and workflow issues that require materials domain knowledge to diagnose.
- Collaborate with AI and data scientists to help shape how experimental data is structured and used for analysis and planning.
You Will Thrive in This Role If You Have
- PhD in Materials Science, Chemistry, Chemical Engineering, or a related field (or equivalent research experience).
- Strong working knowledge of common materials lab hardware (e.g. furnaces, fluid/gas handling manifolds, characterization tools, synthesis equipment) and the workflows built around them.
- Ability to communicate fluently with both scientists and engineers, and to translate scientific intent into clear technical requirements.
- Comfort writing Python to interact with instruments, manipulate experimental data, and prototype automation workflows
Especially Strong Candidates May Also Have
- Prior experience working alongside automation or software engineers.
- Experience with electronic lab notebooks, LIMS, or other lab data systems.
- A track record of designing or adapting experimental protocols for higher-throughput or automated execution.
Mechanics
- Minimum education: PhD or equivalent combination of education and hands-on research experience
- Location: Menlo Park, CA (Soon: San Francisco, too)
- Compensation: $200,000-$250,000 + equity
- Visa sponsorship: Yes, we sponsor visas.
Listed by Periodic Labs for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.