Head of Research, DataLab
Location
Remote
Employment Type
Full time
Location Type
Remote
Department
DataLab
Overview
Application
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
About Protege
We are building Protege to solve the biggest unmet need in AI - getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI's data problem is a generational opportunity. We're backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI - and in tech.
We're a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
DataLab is Protege's research arm - a team of research scientists committed to tackling the fundamental challenges and open questions regarding data for AI. We bridge the gap between research theory and data deployment to push the frontier forward, publishing on the questions that matter: what agentic AI should actually be trained to do, how to quality-control large-scale corpora, and how to build evaluation datasets that reflect the real world rather than the leaderboard.
We're a lean, fast-moving, high-trust team of builders who deeply care about scientific rigor and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
The Role
The Head of Research will lead Protege's DataLab as a strategic research function focused on answering the hardest questions about data for AI. This person will define the research agenda, build rigorous systems for experimentation and evaluation, and ensure DataLab's work directly informs product direction, customer strategy, and platform capabilities.
This is both a leadership and builder role. The right person can set a clear research roadmap, manage and scale a high-performing team, partner across Product, Engineering, GTM, and Leadership, and translate ambiguous technical questions into practical frameworks, measurable experiments, and actionable recommendations.
The Director of Research will work closely with GTM across sales, delivery, and partner conversations — bringing technical credibility into the room when it matters and representing Protege’s research perspective externally.
What You'll Do
Define and lead the research strategy for Protege's Data Lab, aligning experimentation with company priorities and product direction.
Partner closely with Product, Engineering, and GTM teams to identify high-value research opportunities tied to AI data quality, evaluation, and marketplace performance.
Design and oversee experiments that evaluate dataset quality, model performance, synthetic data workflows, and privacy-preserving methodologies.
Build scalable systems for benchmarking, labeling quality analysis, and training data evaluation across multiple AI modalities.
Serve as a customer-facing research partner to GTM, representing DataLab in sales, technical discovery, delivery, and customer strategy conversations while translating research depth into practical guidance that helps Protege deliver value to our partners and customers.
Translate ambiguous technical questions into clear research frameworks, measurable hypotheses, and actionable recommendations.
Publish internal research findings that directly influence product decisions, customer strategy, and platform capabilities.
Lead, manage, and scale a high-performing team of researchers and data scientists, driving execution, technical excellence, and career development across the Data Lab organization.
Establish operational rigor around experimentation, reproducibility, and research documentation.
Represent Protege externally through technical conversations with customers, partners, and the broader AI ecosystem.
Stay at the forefront of advancements in foundation models, evaluation methodologies, data infrastructure, and AI alignment research.
What Success Looks Like
30 Days: Learn and Assess
Build a deep understanding of Protege's platform, customer workflows, and current research priorities.
Meet cross-functional stakeholders across Product, Engineering, GTM, and Leadership to understand strategic goals and existing technical challenges.
Audit current Data Lab workflows, experimentation infrastructure, and evaluation methodologies.
Identify immediate opportunities to improve research velocity, rigor, and cross-functional communication.
60 Days: Establish Direction and Build Momentum
Deliver a clear research roadmap aligned to company priorities and platform differentiation.
Launch initial experiments focused on data quality, evaluation systems, or marketplace optimization.
Introduce standardized processes for experiment tracking, documentation, and reproducibility.
Develop hiring plans and organizational structure for scaling the Data Lab function.
90 Days: Drive Measurable Impact
Deliver research insights that directly influence product roadmap decisions or customer outcomes.
Demonstrate measurable improvements in experimentation speed, evaluation quality, or data performance metrics.
Establish Protege's Data Lab as a trusted strategic partner across Engineering, Product, and GTM.
Recruit, onboard, and effectively lead key technical talent to support long-term research initiatives and company growth.
What we're looking for
You have:
Led impactful research initiatives in AI, machine learning, data infrastructure, or applied research environments where outcomes influenced product or business strategy.
Built and managed high-performing research teams that operated with autonomy, technical rigor, and fast execution cycles.
Experience partnering directly with customers, GTM teams, or external stakeholders in applied technical settings, with the judgment to represent both the research and business value of complex AI/data work clearly and credibly.
Developed frameworks for evaluating model performance, dataset quality, synthetic data, or large-scale experimentation systems.
Experience operating in ambiguous, fast-moving environments where priorities evolved quickly and execution mattered.
Strong ability to communicate complex technical findings to both technical and non-technical audiences.
Demonstrated ownership of cross-functional initiatives spanning research, engineering, product, and go-to-market teams.
Track record of translating research into practical systems, tools, or customer-facing impact.
Experience working with modern AI systems, foundation models, LLM evaluation workflows, or privacy-centric data methodologies.
High judgment, strong prioritization skills, and comfort making decisions with incomplete information.
Motivated by building from first principles and solving foundational problems in AI infrastructure.
Protege's Values
Pass the Loved Ones' Test
We act with integrity and do the right thing - especially when it's hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Protege for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
141 days ago
Binance Accelerator Program - AI Research Scientist (LLM Reasoning & Post-Training)
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. Binance is trusted by more than 320 million people in 100+ countries for its industry-leading security, transparency, trading engine speed, protections for investors, and unmatched portfolio of digital asset products and offerings from trading and finance to education, research, social good, payments, institutional services, and Web3 features. Binance is devoted to building an inclusive crypto ecosystem to increase the freedom of money and financial access for people around the world with crypto as the fundamental means.
About Binance Accelerator Program
Binance Accelerator Program (BAP) is a 3-6 month internship program designed for Early Career talent to have firsthand experience in the rapidly expanding digital assets space. You will be given the opportunity to develop your skills at Binance and understand what it’s like to work at the world's leading blockchain ecosystem. As part of your internship in the BAP, there will also be opportunities for networking and development, which will expand your professional network and build transferable skills to propel you forward in your career. Learn about the BAP Program HERE.
Who may apply
Current university students and recent graduates.
*Terms of employment / engagement shall be subject to contract and local applicable laws
About the Role
You'll work alongside senior research scientists on problems at the frontier of LLM reasoning, post-training methodology, and agentic AI — in one of the few environments where your models interact with live global markets at scale.
This isn't a support or literature-review role. You'll run experiments, form independent hypotheses, implement ideas from recent papers, and work closely with engineering teams to understand how research behaves under real production constraints — 24/7, zero-downtime, hundreds of millions of users.
Who may apply
Current university students (Masters, PHD in AI track) or recent graduates who don't mind starting as intern.
Listed by Binance for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
141 days ago
Binance Accelerator Program - Research Data Scientist
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. Binance is trusted by more than 320 million people in 100+ countries for its industry-leading security, transparency, trading engine speed, protections for investors, and unmatched portfolio of digital asset products and offerings from trading and finance to education, research, social good, payments, institutional services, and Web3 features. Binance is devoted to building an inclusive crypto ecosystem to increase the freedom of money and financial access for people around the world with crypto as the fundamental means.
About Binance Accelerator Program
Binance Accelerator Program (BAP) is a 3-6 month internship program designed for Early Career talent to have firsthand experience in the rapidly expanding digital assets space. You will be given the opportunity to develop your skills at Binance and understand what it’s like to work at the world's leading blockchain ecosystem. As part of your internship in the BAP, there will also be opportunities for networking and development, which will expand your professional network and build transferable skills to propel you forward in your career. Learn about the BAP Program HERE.
Who may apply
Current university students and recent graduates.
*Terms of employment / engagement shall be subject to contract and local applicable laws
About the Role
For all students looking to gain hands-on experience! You'll work directly with senior scientists on research problems at the frontier of LLM reasoning, post-training methodology, and agentic AI — applied to global crypto markets. Your work has a direct path to production systems serving hundreds of millions of users, and where findings warrant it, a clear path to external publication.
You run experiments, implement ideas from recent research, synthesize papers into hypotheses, and work with engineers to understand how research translates to real systems.
This is not a purely literature-review or passive research role. You are expected to think independently and generate insight.
Listed by Binance for a position based in the United States. Employers on this board attest they are hiring domestically.
We are building AI to simulate the world through merging art and science.
We believe that world models are at the frontier of progress in artificial intelligence. Language models alone won’t solve the world’s hardest problems – robotics, disease, scientific discovery. Real progress requires models that experience the world and learn from their mistakes, the same way that humans do. And this kind of trial and error can be massively accelerated when done in simulation, rather than in the real world.
World models offer the most clear path to general-purpose simulation, changing how stories are told, how scientific progress is made and how the next frontiers of humanity are reached.
Our team consists of creative, open minded, caring and ambitious people who are determined to change the world. We aspire to continuously build impossible things and our ability to do so relies on building an incredible team. If you are driven to do the same, we'd love to hear from you.
ABOUT THE ROLE
*Open to hiring remote across the US — preferable to have someone near an office in NYC or San Francisco.
This is a rare PM role with true vertical ownership across the full product lifecycle; from early-stage research and pre-training through to post-training and UX buildout through to GTM strategy and pricing - giving the right person a GM-like mandate to drive product decisions end-to-end in a way that's unique for a company of our scale.
We’re looking for a passionate Product Manager to help us develop cutting-edge AI tools. This role is for someone who has a combination of strong technical or ML background and experience in product management. You have the ability to take the state-of-the-art AI systems our research team builds and turn them into impactful tools for our customers.
You’ll deeply understand what users want, what’s possible to build, and what’s strategic for our business. You’ll be a bridge between our users and our internal teams so we ship products that have the most value for Runway’s creators.
This role is a fundamental and critical part of the team for a market-defining company in its early stages.
WHAT YOU’LL DO
- Work alongside our research team to define new AI applications that solve customer needs and drive business value
- Collaborate closely with engineering and design to productize research and continuously improve existing tools
- Define and analyze metrics, talk with users weekly, and conduct A/B tests to discover use cases, validate pain points, and prioritize features
WHAT YOU’LL NEED
- 1+ years of direct experience with SQL, python, code generation tools, and version control best practices
- 4+ years of product management experience
- Familiarity with the current state-of-the-art ML capabilities and ML products
- Experience working with fast-moving teams and building products from scratch
- Passion for seeing research through from initial conception to ultimate application
- Excellent written and verbal communication skills
Runway strives to recruit and retain exceptional talent from diverse backgrounds while ensuring pay equity for our team. Our salary ranges are based on competitive market rates for our size, stage and industry, and salary is just one part of the overall compensation package we provide.
There are many factors that go into salary determinations, including relevant experience, skill level and qualifications assessed during the interview process, and maintaining internal equity with peers on the team. The range shared below is a general expectation for the function as posted, but we are also open to considering candidates who may be more or less experienced than outlined in the job description. In this case, we will communicate any updates in the expected salary range.
Lastly, the provided range is the expected salary for candidates in the U.S. Outside of those regions, there may be a change in the range, which again, will be communicated to candidates.
WORKING AT RUNWAY
Great things come from great teams. https://www.youtube.com/watch?v=kwmj4ato2kw&ab_channel=Runway We’d love to hear from you.
We’re committed to creating a space where our employees can bring their full selves to work and have equal opportunity to succeed. So regardless of race, gender identity or expression, sexual orientation, religion, origin, ability, age, veteran status, if joining this mission speaks to you, we encourage you to apply.
More about Runway
- Universal World Simulator https://runwayml.com/world-simulator.html
- GWM-1 https://runwayml.com/research/introducing-runway-gwm-1
- Gen-4.5 https://runwayml.com/research/introducing-runway-gen-4.5
- General World Models https://runwayml.com/research/introducing-general-world-models
- Robotics SDK https://runwayml.com/research/introducing-runway-gwm-1#robotics-section
- Conversational Real-time Agents https://runwayml.com/research/introducing-runway-gwm-1#avatars-section
- Runway Studios https://runwayml.com/studios
We're excited to be recognized as a best place to work:
Crain's https://www.crainsnewyork.com/awards/best-places-work-2023 | InHerSight https://www.inhersight.com/companies/best/city/new-york-city-ny | BuiltIn NYC https://builtin.com/awards/new-york-city/2024/best-places-to-work | INC https://www.inc.com/best-workplaces/2024
Listed by RunwayML for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
168 days ago
Senior Director, Clinical and Translational Research
Back to jobs
Senior Director, Clinical and Translational Research
Remote, USA
Apply
Omada Health is on a mission to bend the curve of chronic disease.
Job Overview:
The Senior Director, Clinical & Translational Research leads Omada’s efforts to demonstrate the clinical value economic and population health value of its programs across various populations.. This role collaborates closely with Clinical Strategy and Commercial teams to generate compelling clinical and health economics evidence and translate findings into the market. This role also collaborates closely with Clinical Innovation and Product efforts to personalize and optimize Omada’s program impact in service of improving member engagement and outcomes.
Core Responsibilities:
Clinical Research & Commercialization
Design, oversee, and execute clinical research studies and health economics research studies and analyses, including randomized controlled trials, observational cohort studies, economic models, healthcare claims analyses, and real-world evidence (RWE), evaluating the clinical and economic outcomes and highlighting the value associated with Omada’s digital health programs.
Lead the development and interpretation of health economics and outcomes research (HEOR), including cost-effectiveness, budget impact, and total cost of care analyses that clearly articulate value to Medicare Advantage plans, commercial employers, and provider partners.
Lead and mentor a team of PhD-level scientists (4) and project manager, providing guidance, coaching, and professional development opportunities.
Manage external research vendors, consultants, and partners.
Ensure compliance with regulatory requirements, ethical standards, and industry best practices in all research activities and publications.
Publish research findings in peer-reviewed journals and white papers.
Author Thought Leadership pieces to represent the company’s position on relevant clinical topics.
Collaborate closely with Clinical Strategy and Commercial teams, including Sales and Marketing, to develop compelling and evidence-based value stories and to translate research findings to Commercial-facing assets.
Program Development and Experimentation
Collaborate with Clinical Innovation (including Product, Care Delivery, and Data) in the development of personalized program experiments and quality improvement projects in service of improving member engagement and outcomes.
Ensure sound and ethical scientific practices are followed, clinical rigor is maintained, and experiments are grounded in evidence-based behavioral science with high likelihood of making a meaningful impact on engagement and clinical outcomes, quality metrics and economic value.
Seek broad input from cross-functional teams (e.g., Marketing/Sales, Clinical, Care Team) to determine business impactful and clinically meaningful metrics for key program development pilots. Generate and oversee the Evaluation Plans for the pilots to ensure they are designed to meet business impactful metrics.
Collaborate with the Subject Matter Experts (SME) for Behavior Change and Behavioral Health with R&D to inform program personalization and clinical impact, including in the development of meaningful interventions for programmatic pilots.
Serve as a Subject Matter Expert (SME) for Clinical Outcomes and Health Economics with Commercial Teams, including internal consultation as well as customer-facing meetings with consultants, actuaries and data-savvy customers.
Qualifications:
Advanced degree (Ph.D., Pharm.D., or M.D.) in health services research, psychology, public health, or related fields
Minimum of 10 years of relevant clinical research experience
Deep expertise in health economics and outcomes research (HEOR), including hands-on and/or leadership experience with healthcare claims analysis, cost and utilization modeling, and total cost of care analyses for Medicare Advantage and commercial employer populations.
Familiarity with quality measurement frameworks such as HEDIS and Medicare Advantage Star Ratings, including experience using these measures to evaluate and communicate program impact.
Demonstrated track record of designing and leading studies, including clinical trials, observational studies, economic evaluations, and patient-reported outcomes research.
Strong leadership and team management skills, with the ability to inspire and motivate cross-functional teams towards common goals.
Exceptionally collaborative and adaptable, with the ability to forge strong relationships across multiple functions and domains and to flex to meet business priorities.
Excellent communication and presentation skills, with the ability to effectively convey complex scientific concepts to diverse audiences.
Strategic thinker with a results-oriented mindset and the ability to thrive in a fast-paced, entrepreneurial environment.
Ability and readiness for quick growth opportunities. This is a fast track position with potential for VP level responsibilities within 1 year.
Benefits:
Competitive salary with generous annual cash bonus
Equity grants
Employee stock purchasing plan (ESPP)
Remote first work from home culture
Flexible Time Off to help you rest, recharge, and connect with loved ones
Generous parental leave
Health, dental, and vision insurance (and above market employer contributions)
401k retirement savings plan
Lifestyle Spending Account (LSA)
Mental Health Support Solutions
...and more!
It takes a village to change health care. As we build together toward our mission, we strive to embody the following values in our day-to-day work. We hope these hold meaning for you as well as you consider Omada!
Cultivate Trust. We listen closely and we operate with kindness. We provide respectful and candid feedback to each other.
Seek Context. We ask to understand and we build connections. We do our research up front to move faster down the road.
Act Boldly. We innovate daily to solve problems, improve processes, and find new opportunities for our members and customers.
Deliver Results. We reward impact above output. We set a high bar, we’re not afraid to fail, and we take pride in our work.
Succeed Together. We prioritize Omada’s progress above team or individual. We have fun as we get stuff done, and we celebrate together.
Remember Why We’re Here. We push through the challenges of changing health care because we know the destination is worth it.
About Omada Health: Omada Health (Nasdaq: OMDA) is reverse engineering the way healthcare is delivered in America, putting the space between doctor visits–where health is won or lost–at the center of care. Today's healthcare system poorly serves chronic conditions that require ongoing support outside of the exam room, like obesity, diabetes, hypertension, cholesterol, and musculoskeletal conditions. Omada’s virtual-first model combines human-led care teams, connected devices, and AI-enabled technology to deliver personalized care at scale, including support for GLP-1 therapy. Omada has served more than two million members since launch across 2,000+ employers, health plans, pharmacy benefit managers, and health systems. Learn more at omadahealth.com.
Omada is thrilled to share that we’ve been certified as a Great Place to Work! Please click here for more information.
We carefully hire the best talent we can find, which means actively seeking diversity of beliefs, backgrounds, education, and ways of thinking. We strive to build an inclusive culture where differences are celebrated and leveraged to inform better design and business decisions. Omada is proud to be an equal opportunity workplace and affirmative action employer. We are committed to equal opportunity regardless of race, color, religion, sex, gender identity, national origin, ancestry, citizenship, age, physical or mental disability, legally protected medical condition, family care status, military or veteran status, marital status, domestic partner status, sexual orientation, or any other basis protected by local, state, or federal laws. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Below is a summary of salary ranges, by geographic zone, for this role*. These ranges represent a good faith estimate of the minimum and maximum base salary for the position in that zone at the time of posting. Please refer to this resource for detail on zone designations.
Zone 1: $243,800 - $304,800
Zone 2: $233,200 - $291,500
Zone 3: $212,000 - $265,000
Please note that zones may be updated as market data changes but we will honor the range applicable to your zone (which is determined by your location) as of your application date.
*The actual offer, including the compensation package, will be within the posted range for the relevant zone and will be determined based on multiple factors, such as the candidate's skills and experience, and other business considerations such as internal equity.
Please click here for more information on our Candidate Privacy Notice.
Apply for this job
*
indicates a required field
Autofill my application
First Name*
Last Name*
Email*
Phone
Country*
Phone*
Location (City)*
Locate me
Resume/CV*
Attach
Attach
Dropbox
Google Drive
Enter manually
Enter manually
Accepted file types: pdf, doc, docx, txt, rtf
Cover Letter
Attach
Attach
Dropbox
Google Drive
Enter manually
Enter manually
Accepted file types: pdf, doc, docx, txt, rtf
LinkedIn Profile
Website
Do you currently live in the United States?
Select...
We are unable to employ candidates residing outside of the US. Due to the nature of our work, data, specifically personal health information may not be accessed, disclosed or used outside of the U.S. Please note that the U.S. is limited to the 50 states of the United States.
Which best describes your highest degree and years of clinical/health services research experience? *
Select...
Which types of HEOR work have you personally led? Select all that apply. *
Healthcare claims analysis
Cost and utilization modeling
Total cost of care analysis
Other HEOR (e.g., budget impact, burden of illness)
I have not led HEOR work
What is the highest level of responsibility you have had for research studies?*
Select...
In which areas have you led or co-led research? Select all that apply.*
Select...
What best describes your experience managing research teams?*
Select...
Voluntary Self Identification
At Omada, we’re working every day to build a team where people from all backgrounds can thrive. To help us learn more about how we can increase diversity in our candidate pool, we invite you to provide demographic information in a confidential survey. Your responses will not be associated with your specific application and will not in any way be used in the hiring decision.
I identify my gender as:
Select...
I identify as transgender:
Select...
I identify my sexual orientation as:
Select...
Which of the following best represents your racial or ethnic background: (Check more than one if that is how you identify)
Select...
Veteran Status:
Select...
Do you identify as a person with a disability of any kind, visible or invisible?
Select...
Submit application
Powered by
Listed by Omada Health for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
183 days ago
Manager, Quality – Cell Therapy Site Onboarding (West Coast)
More than one million people in the United States today are fighting blood cancer. While a traditional allogeneic stem cell transplant has been the best hope for many, the transplant itself can prove fatal or lead to serious conditions, such as graft vs. host disease. Orca Bio is a commercial-stage biotechnology company redefining the transplant process by developing next-generation cell therapies with the goal of providing significantly better survival rates with dramatically fewer risks.With our purified, high-precision cell therapies we hope to not only replace patients' blood and immune systems with healthy ones, but also restore their lives.
Position Summary: The Manager, Quality – Commercial Site Onboarding will oversee and manage the quality component of the onboarding process for our commercial sites (Authorized Treatment Centers) on the West Coast. This role will be responsible for ensuring a smooth, efficient, and high-quality onboarding experience, ensuring that all commercial sites are prepared for successful implementation and ongoing operations. The ideal candidate will have a strong background in project management, quality management or customer relationship management, ideally within biotech/bone marrow transplant/cell therapies, with a focus on delivering exceptional service and results.
Travel: up to 25-50% travel (to assigned Authorized Treatment Centers) will be required. Candidates must reside on the West Coast.
Listed by Orca Bio for a position based in the United States. Employers on this board attest they are hiring domestically.
ABOUT POOLSIDE
In this decade, the world will create Artificial General Intelligence. There will only be a small number of companies who will achieve this. Their ability to stack advantages and pull ahead will define the winners. These companies will move faster than anyone else. They will attract the world's most capable talent. They will be on the forefront of applied research, engineering, infrastructure and deployment at scale. They will continue to scale their training to larger & more capable models. They will be given the right to raise large amounts of capital along their journey to enable this. They will create powerful economic engines. They will obsess over the success of their users and customers.
Poolside exists to be this company: to build a world where AI will be the engine behind economically valuable work and scientific progress. We believe the fastest way to reach AGI lies in accelerating software development itself, by reshaping the developer experience with agentic systems, coding assistants, and the frontier models that power them. We deploy these systems directly into the development environments of security-conscious enterprises.
ABOUT OUR TEAM
We were founded in the US and have our home there, but our team is distributed across Europe and North America. We get our fix of in-person collaboration in Paris each month for 3 days, with an open invitation to stay the whole week. For those based in PST, we understand this is a significant travel cadence; we are open to agree on a lower cadence and will discuss this in the interview process. We also do longer off-sites once a year.
Our team is a multidisciplinary blend of research, engineering, and business experts. What unites us is our deep care for what we build together. We’re in a race that requires hard work, intellectual curiosity, and obsession; to balance this intensity, we’ve assembled a team of low ego and kind-hearted individuals who have built the special culture Poolside has. By building collaboratively and with intention, we create a compounding effect that moves the entire company forward towards our mission: reaching AGI through intelligence systems built for software development.
ABOUT THE ROLE
You’ll be working on our data team focused on the quality of the datasets being delivered for training our models. This is a hands-on role where your #1 mission would be to improve the quality of our datasets across the entire training cycle (pre-training, mid-training, post-training, RL) by leveraging your previous experience, intuition and training experiments. This role particularly focuses on generating synthetic data at scale and determining the best strategies to leverage such data into training large models. You’ll closely collaborate with other teams like Pre-training, Pre-training data, Post-training, RL2L, Evals, and Product to define high-quality data needs that map to missing model capabilities and downstream use cases.
Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role. You will constantly lead original research initiatives through short, time-bounded experiments while deploying highly technical engineering solutions into production. With the volumes of data to process being massive, you'll have a performant distributed data pipeline together with large GPU clusters at your disposal.
Curious about the tech? Take a deep dive into our data work in our Laguna M.1/XS.2 Technical Report. https://arxiv.org/abs/2605.27605
YOUR MISSION
To deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class Poolside coding agents.
RESPONSIBILITIES
- Follow the latest research related to LLMs and synthetic data generation in particular. Be familiar with the most relevant open-source datasets and models.
- Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing the resources available.
- Closely with cross-team to ensure the experiments run and data generated is the most efficient use of compute and time resources for the improvements in quality of our models.
- Continuously measure and refine the quality of the datasets being generated while validating the final data strategy through quantitative data ablation experiments.
SKILLS & EXPERIENCE
- Strong machine learning and engineering background
- Experience with Large Language Models (LLM), including:
- Understanding of how LLMs learn
- Data ablations and scaling laws
- Post-training techniques
- Training reasoning and agentic models
- Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc.
- Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc)
- Experience in building trillion-scale pretraining datasets, and familiarity with concepts like data curation, deduplication, data mixing, tokenization, curriculum, impact of data repetition, etc.
- Excellent programming skills in Python
- Strong prompt engineering skills
- Experience working with large-scale GPU clusters and distributed data pipelines
- Strong obsession with data quality
- Research experience:
- Nice to have: Author of scientific papers on any of the topics: applied deep learning, LLMs, source code generation, etc.
- Nice to have: Formal machine learning, mathematics, or computer science background
- Can freely discuss the latest papers and descend to fine details
- Is reasonably opinionated
PROCESS
- Intro call with one of our Founding Engineers
- Technical Interview(s) with one of our Members of Engineering
- Team fit call with the People team
- Final interview with one of our Founding Engineers
BENEFITS
- Fully remote work & flexible hours
- 37 days/year of vacation & holidays
- Health insurance allowance for you & dependents
- 16 weeks of flexible, full-pay parental leave
- Company-provided equipment
- Well-being, always-be-learning & home office allowances
- Frequent team get togethers
- Diverse & inclusive people-first culture
Listed by poolside for a position based in the United States. Employers on this board attest they are hiring domestically.
ABOUT POOLSIDE
In this decade, the world will create Artificial General Intelligence. There will only be a small number of companies who will achieve this. Their ability to stack advantages and pull ahead will define the winners. These companies will move faster than anyone else. They will attract the world's most capable talent. They will be on the forefront of applied research, engineering, infrastructure and deployment at scale. They will continue to scale their training to larger & more capable models. They will be given the right to raise large amounts of capital along their journey to enable this. They will create powerful economic engines. They will obsess over the success of their users and customers.
Poolside exists to be this company: to build a world where AI will be the engine behind economically valuable work and scientific progress. We believe the fastest way to reach AGI lies in accelerating software development itself, by reshaping the developer experience with agentic systems, coding assistants, and the frontier models that power them. We deploy these systems directly into the development environments of security-conscious enterprises.
ABOUT OUR TEAM
We were founded in the US and have our home there, but our team is distributed across Europe and North America. We get our fix of in-person collaboration in Paris each month for 3 days, with an open invitation to stay the whole week. For those based in PST, we understand this is a significant travel cadence; we are open to agree on a lower cadence and will discuss this in the interview process. We also do longer off-sites once a year.
Our team is a multidisciplinary blend of research, engineering, and business experts. What unites us is our deep care for what we build together. We’re in a race that requires hard work, intellectual curiosity, and obsession; to balance this intensity, we’ve assembled a team of low ego and kind-hearted individuals who have built the special culture Poolside has. By building collaboratively and with intention, we create a compounding effect that moves the entire company forward towards our mission: reaching AGI through intelligence systems built for software development.
ABOUT THE ROLE
You’ll be working on our data team focused on the quality of the datasets being delivered for training our models. This is a hands-on role where your #1 mission would be to improve the quality of our datasets across the entire training cycle (pre-training, mid-training, post-training, RL) by leveraging your previous experience, intuition and training experiments. This role particularly focuses on generating synthetic data at scale and determining the best strategies to leverage such data into training large models. You’ll closely collaborate with other teams like Pre-training, Pre-training data, Post-training, RL2L, Evals, and Product to define high-quality data needs that map to missing model capabilities and downstream use cases.
Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role. You will constantly lead original research initiatives through short, time-bounded experiments while deploying highly technical engineering solutions into production. With the volumes of data to process being massive, you'll have a performant distributed data pipeline together with large GPU clusters at your disposal.
Curious about the tech? Take a deep dive into our data work in our Laguna M.1/XS.2 Technical Report. https://arxiv.org/abs/2605.27605
YOUR MISSION
To deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class Poolside coding agents.
RESPONSIBILITIES
- Follow the latest research related to LLMs and synthetic data generation in particular. Be familiar with the most relevant open-source datasets and models.
- Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing the resources available.
- Closely with cross-team to ensure the experiments run and data generated is the most efficient use of compute and time resources for the improvements in quality of our models.
- Continuously measure and refine the quality of the datasets being generated while validating the final data strategy through quantitative data ablation experiments.
SKILLS & EXPERIENCE
- Strong machine learning and engineering background
- Experience with Large Language Models (LLM), including:
- Understanding of how LLMs learn
- Data ablations and scaling laws
- Post-training techniques
- Training reasoning and agentic models
- Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc.
- Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc)
- Experience in building trillion-scale pretraining datasets, and familiarity with concepts like data curation, deduplication, data mixing, tokenization, curriculum, impact of data repetition, etc.
- Excellent programming skills in Python
- Strong prompt engineering skills
- Experience working with large-scale GPU clusters and distributed data pipelines
- Strong obsession with data quality
- Research experience:
- Nice to have: Author of scientific papers on any of the topics: applied deep learning, LLMs, source code generation, etc.
- Nice to have: Formal machine learning, mathematics, or computer science background
- Can freely discuss the latest papers and descend to fine details
- Is reasonably opinionated
PROCESS
- Intro call with one of our Founding Engineers
- Technical Interview(s) with one of our Members of Engineering
- Team fit call with the People team
- Final interview with one of our Founding Engineers
BENEFITS
- Fully remote work & flexible hours
- 37 days/year of vacation & holidays
- Health insurance allowance for you & dependents
- 16 weeks of flexible, full-pay parental leave
- Company-provided equipment
- Well-being, always-be-learning & home office allowances
- Frequent team get togethers
- Diverse & inclusive people-first culture
Listed by poolside for a position based in the United States. Employers on this board attest they are hiring domestically.