Applied Healthcare Researcher
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
Role Overview
We are hiring Applied Healthcare Researchers to join a team within DataLab focused entirely on healthcare training data.
Our customers are researchers at the frontier labs and AI startups building specialized healthcare models. They come to us with model-development problems, not dataset specifications. Figuring out which healthcare data actually solves their problem, and proving that it does, is the research question we answer in DataLab.
In this role you will work directly with researchers at those labs to understand what they're trying to train or evaluate, determine what healthcare data can support it, and do the research needed to demonstrate that it will. This is fast-iterating, customer-facing research on a customer's timeline. You will be the primary technical and research link to the customer — not a technical resource brought in for credibility, but the person driving the conversation and pulling in the solutions, engineering, and data partnerships teams as needed.
Core Responsibilities
Customer Research Partnership
You will be the research partner to AI researchers at frontier labs and startups who are working on healthcare problems.
• Serve as the primary technical and research point of contact for healthcare customer conversations.
• Translate a lab's model-development goals into concrete, feasible data strategies.
• Help customers scope opportunities and identify the highest-value data available to them.
• Explain data limitations, tradeoffs, and potential biases to technically sophisticated stakeholders while grounding conversations in what real-world data actually looks like.
• After delivery, answer the research questions customers raise about the data we provided. Delivery is not the end of the relationship.
Applied Research & Method Development
Curating the right data product is a research problem, and you'll own solving it.
• Develop and evaluate methods — fine-tuning, LLM-based extraction, classification, rules-based approaches, or whatever the problem calls for — to demonstrate that a dataset can support a customer's training or evaluation objective.
• Design and run feasibility research pre-contract: can this data support this model objective, at what quality, with what caveats.
• Build the evidence base that makes a data strategy credible — benchmarks, validation analyses, error characterization, and honest assessments of where the data falls short.
• Partner with the Assessments team on healthcare benchmarks across modalities.
Data Feasibility & Dataset Strategy
• Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data.
• Identify proxy variables or alternative dataset structures when the ideal variable doesn't exist.
• Analyze partner and source datasets — schema, field availability, quality, completeness, and required transformations.
• Contribute to our point of view on which healthcare data matters most for which modality and which stage of model development.
• Help evaluate new data partners and identify datasets worth acquiring before a customer asks for them.
Reusable Research & Scaling
• Produce reusable research, evidence, and technical collateral rather than starting from scratch for each opportunity.
• Identify where a successful one-off approach should become a repeatable workflow, and work with Product and Engineering to operationalize it.
• Help expand proven healthcare datasets across multiple customers instead of selling them once.
Cross-Functional Collaboration
• Work with Solutions and FDEs from the beginning of an opportunity.
• Coordinate with Healthcare Data Partnerships on sourcing and with Product and Engineering on tooling.
Required Experience & Skills
• Advanced degree (PhD or Master's plus 2+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience.
• Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data.
• Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training.
• Strong Python and SQL, with the ability to work independently against large datasets.
• Experience designing evaluations — measuring data quality and dataset representativeness.
• Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans.
• Comfort operating on a customer's timeline without lowering the standard of the research.
Ideal Profile
The ideal candidate:
• Is energized by working directly with customers, and specifically by working with other researchers as peers.
• Moves fast on messy, real-world problems and knows which corners can and cannot be cut.
• Is rigorous about what the data can and cannot support, and willing to tell a customer when the answer is no.
• Enjoys the full arc — scoping a vague problem, doing the research, and showing the result to the person who asked for it.
• Thinks about leverage: builds the reusable version rather than the one-off when it's worth doing.
About DataLab
DataLab exists because truly useful data is rare — and the frontier of AI development only moves forward when high-quality data makes it possible.
We believe data is one of the most underdeveloped layers of the AI stack. Our work focuses on building and evaluating high-value datasets grounded in real-world workflows and economically meaningful tasks. Our research spans data quality, evaluation design, privacy-preserving transformation, and task-grounded AI training data.
Protege Values
Pass the Loved Ones’ Test
We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Forward Deployed Machine Learning Engineer
Full Stack Software Engineer
Strategic Account Manager
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
About the Role
The Strategic Account Manager owns one of the most consequential commercial relationships at Protege: th Foundation labs and Frontier AI builders — organizations like OpenAI, Meta, Google, Apple, and Anthropic — that are defining what AI becomes.
This is not a conventional sales role. A SAM's job is not to simply to close a series of unrelated deals. A SAM's job is to engineer high-value relationships and drive unimpeachable value, en route to driving significant GMV for Protege. You are the primary commercial owner of a small number of accounts that represent tens of millions of dollars of annual revenue potential — and the relationships that make Protege's broader business possible.
You will work in close partnership with DataLab researchers, Product Success, and Partnerships to translate complex, multi-modal data needs into real solutions — often before our customers fully know what they need. You understand how foundation models are built, what researchers care about, and how to show up as a trusted partner rather than a vendor. You think in multi-year relationship arcs, not quarterly close dates.
What You'll Do
• Own a portfolio of 2–4 foundation lab or frontier AI accounts end-to-end — from relationship development and deal structuring through contracting, delivery handoff, and expansion.
• Map the full account landscape: identify key researchers, data science leads, procurement stakeholders, and executive sponsors. Understand who influences decisions and who makes them.
• Lead business-level discovery across departments and functions to identify projects and opportunities for Protege to broaden and deepen our penetration of each account
• Coordinate the teams who interact technically with our customer research teams -You’ll work closely with our Solutions and DataLab teams to do so
• Build and maintain a living account plan for each strategic account, including deal chronology, relationship health, 12-month GMV potential, open opportunities, and known risks.
• Orchestrate internal resources — DataLab, Solutions, Partnerships, Legal, Customer Success Manger (CSM) — to deliver on commitments and protect relationship equity. You are the single point of commercial accountability.
• Drive multi-year, multi-modal deal strategy: structure agreements that deepen account relationships over time rather than optimizing for individual transaction size.
• Surface patterns across your accounts that inform Protege's data supply strategy, product roadmap, and go-to-market — you are on the frontier of what the world's best AI labs need next.
• Maintain rigorous deal hygiene: accurate forecasting, clean Deal Desk submissions, documented handoffs to CSM and Product Success at close.
• Represent Protege at field events, research dinners, and conferences where foundation lab relationships are built and deepened.
What Success Looks Like
30 Days — Learn the accounts. Earn the right to lead them.
• Complete deep account briefings on each assigned account: deal history, active relationships, open opportunities, delivery status, and known risks.
• Shadow existing deal conversations, Deal Desk sessions, and DataLab research calls to understand how Protege's internal machine works.
• Establish direct relationships with your primary points of contact at each account — researchers, data leads, and commercial stakeholders.
• Identify the single highest-priority open opportunity in each account and develop a clear point of view on what it will take to close it.
60 Days — Lead the commercial motion. Deepen the relationships.
• Own all commercial conversations in your accounts, with DataLab and Product Success running alongside you — not in front of you.
• Have a credible 90-day pipeline view for each account: committed deals, active opportunities, and what is needed to progress each.
• Complete account plans for each strategic account in the standard format, reviewed with Don and the exec team.
• Identify at least one expansion opportunity per account that did not exist before your arrival — a new data type, a new team, a new use case.
90 Days — Drive GMV. Build the multi-year relationship arc.
• Close or advance at least one material deal in each account. Material means it moves the quarterly number and deepens the relationship.
• Deliver a 90-day account retro for each strategic account: what we learned, what the relationship looks like now versus on day one, what the 12-month opportunity is, and what we need internally to capture it.
• Have a clear point of view on the path to multi-year agreements in at least one account — what it would take, what the unlock is, and when.
• Be the person DataLab researchers want to bring into a room and the person customers ask for by name.
What You Bring
We focus on what you have accomplished, not how long you were in a particular role.
• Proven track record managing and growing large, complex enterprise accounts — preferably with foundation labs, hyperscalers, or frontier AI organizations. You have closed multi-million dollar deals and know what it takes to sustain them.
• Technical fluency sufficient to earn the trust of AI researchers and data scientists. You do not need to be an ML engineer, but you need to understand how models are trained, why data quality and provenance matter, and what de-identification means in practice.
• Demonstrated ability to build relationships at multiple levels of a single organization simultaneously — from working-level researchers to VP and executive sponsors — and to advance deals through complex, multi-stakeholder procurement processes.
• Strong deal architecture skills: you know how to structure agreements that create long-term value, not just close quarters. Experience with multi-year licensing, data access frameworks, or platform agreements is a genuine plus.
• Operational discipline: you run clean pipeline, forecast accurately, and hand off to delivery teams in ways that set them up to succeed rather than scramble.
• High emotional intelligence and judgment in relationship-sensitive situations. Foundation lab accounts require a long game — you know when to push and when to listen, and you never sacrifice trust for a transaction.
• Comfort with ambiguity and fast-moving deal environments. Scope changes, new requirements, and compressed timelines are the norm. You stay organized and calm when others don't.
Protege Values
Pass the Loved Ones’ Test
We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Senior Full Stack Software Engineer
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
About the Role
Protege is hiring a Senior Software Engineer for our Presentation team, the squad that builds the surfaces where Protege's data catalog meets the people who use it. That means two products: the internal dashboard our team uses to catalog, curate, package datasets, and the customer portal where AI labs browse, stream, and evaluate content before they license it.
This is a product engineering role in a TypeScript and Next.js codebase. You'll talk directly with the people who use what you build, design the interfaces yourself, and ship the whole feature: UI, API, data model, and tests. The datasets behind these screens are large and multimodal (video, audio, motion capture, documents, medical imaging), so a lot of the craft is making that volume feel fast, legible, and good to use in a browser.
As a senior engineer, you'll take on the largest and most ambiguous initiatives on the team: the ones that span both apps, cross into other squads' services, or need someone to figure out what the right product even is before building it. You'll partner with the team lead on direction and priorities, lead design on complex workflows, and raise the bar for everyone shipping in this codebase. This role is ideal for product engineers who want to design what they build, talk to the people who use it, and own features end to end.
What You’ll Do:
• Lead the team's larger initiatives from discovery through design, build, and rollout, such as search previews at catalog scale or the presentation of a new content modality
• Talk directly with the partnerships, solutions, and sales teams who use what you build; understand the business context behind each request and bring that view back to the team's priorities
• Design complex workflows and data-heavy interfaces yourself: information hierarchy, interaction patterns, and every state a screen can be in
• Build the workflows our partnerships and solutions teams use every day: sample creation, customer access management, and delivery exports
• Build the customer-facing experience for browsing, streaming, and evaluating multimodal samples
• Make large multimodal datasets feel instant in a browser: streaming playback, previews, lazy loading, and export at scale
• Build UI on top of embedding-based search so users can find the right clips within catalogs of millions
• Decide what runs client-side, in a route handler, or in a backend service, and work with our data teams when the answer crosses squad boundaries
• Handle sensitive data correctly: scoped access, secure share links, and PHI-aware routes
• Set the bar for quality on the squad: testing, performance, design consistency, and UX detail
• Lead design discussions, review code, and unblock other engineers
• Turn repeated one-off UI and workflow patterns into reusable components in our shared design system
What Success Looks Like:
30 days: Learn and Build Relationships
• Get productive in the codebase and ship your first improvements to both apps
• Build a working map of the presentation stack: the two apps, the services behind them, and how content flows from ingestion to a customer's screen
• Meet the partnerships, solutions, and data teams; understand how the catalog gets licensed and delivered, and where the current tooling slows a deal down
60 days: Develop and Cultivate
• Take a major feature or workflow from ambiguous ask to shipped, including its design
• Start raising the bar on quality, design consistency, and performance across the squad's code
• Become the engineer others come to when they're stuck in the presentation layer
90 days: Make an Impact
• Lead a significant initiative end to end, from talking to the people who need it through design, build, and rollout
• Identify at least one leverage opportunity (a reusable component, an architectural improvement, a workflow that should be self-serve) and drive it
• Have a visible effect on how the team ships: patterns adopted, reviews that teach, fewer things falling through cracks
What You Bring:
Must Haves
• 5+ years building production web applications that real users depend on
• Deep TypeScript and React experience, with Next.js or a comparable full-stack framework in production
• You've shipped user-facing features end to end: UI, API, data model, and tests
• Strong design judgment: you can design a complex workflow yourself, hold a bar for visual and interaction detail, and articulate why one UI decision beats another
• Solid backend fundamentals: API design, Postgres, auth and permission modeling, and cloud services (we use AWS)
• You've built data-heavy or media-heavy interfaces and know how to keep them fast: streaming, pagination, caching, and knowing what not to load
• You've worked directly with the users of what you build and let that shape what you shipped
• Attention to detail without losing speed, and a bias to action
• Curious and proactive
Nice to Haves
• Experience with media playback in the browser (HTML5 video/audio, HLS, containers and codecs)
• Experience building UX over embeddings or vector search
• A portfolio, side projects, or shipped work that shows your design taste
• Experience building or maintaining a design system across multiple apps
• Familiarity with auth systems (we use Clerk) and permission modeling
• Experience working with sensitive or regulated data (HIPAA, PHI)
• Prior startup experience as a founding or early engineer
Protege Values
Pass the Loved Ones’ Test
We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Financial Operations Associate
Technical Product Manager, Data Ingestion & Quality
Senior Software Engineer, Data Processing
Forward Deployed Engineer, New Verticals
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
Role Overview
We’re hiring a founding Forward Deployed Engineer to help build a new vertical from the ground up.
You’ll be the first FDE dedicated to this vertical, working directly with the GM to define the market, strategy, and early commercial motion. Your job is to turn early customer demand into durable technical capability: defining what the vertical needs, building reusable infrastructure on top of our existing platform, and establishing the technical patterns that future engagements and future hires can build on.
This is not a standard implementation role. It sits at the intersection of engineering, product judgment, and customer reality. You should be excited to work from first principles, operate in ambiguity, iterate quickly, and partner with product engineering to make strong calls about what should become a core platform capability versus what should remain vertical-specific.
What You'll Do
Build the Technical Foundation
Partner with the GM and early customers to define what the vertical actually needs technically.
Build the first MVP of reusable patterns, integrations, and tooling that future engagements will run on, leveraging our core platform where it fits and extending it where the vertical requires something new.
Make architectural decisions about what belongs in the platform layer versus vertical-specific tooling.
Create the initial technical playbook so future FDEs can build on a real foundation rather than starting from scratch.
Own First Deals End-to-End
Lead the first customer engagements in the vertical, from technical scoping through delivery and post-launch support.
Write robust code that solves immediate customer problems while compounding into reusable infrastructure.
Navigate real-world complexity across customer data, integrations, workflows, and stakeholder dynamics.
Translate messy customer requirements into systems that are durable and maintainable.
Shape What Becomes Product
Partner with Product and Engineering to identify which patterns from early customer work should become core platform capabilities.
Surface repeatable use cases, infrastructure gaps, and product opportunities from live engagements.
Help determine when the vertical is ready to evolve from bespoke delivery into a repeatable product motion.
Partner Across the Company
Work directly with the GM or Solutions Lead to define the vertical’s technical strategy and commercial approach.
Partner with Data Lab on domain-specific data and research questions.
Collaborate with other FDEs on shared patterns, tools, and approaches that should compound across verticals.
Serve as the technical voice of the vertical as it grows.
What Success Looks Like
The shape of this role depends on the vertical you're deployed into, so we measure success based on trajectory rather than a fixed set of outputs. In the first 90 days, we expect the following to happen:
Build an understanding of the vertical and strategy
Develop a strong understanding of the vertical, the market dynamics, and the GM's strategy for building and scaling the business. Gain context on the customer landscape, commercial motion, and the unique technical requirements that will shape the vertical's success.
Understand customer and platform needs
Build a deep understanding of what customers and data partners need, what Protege's platform can already support, and where meaningful gaps exist. Develop a clear point of view on which constraints are temporary, which require new infrastructure, and which represent opportunities for future product investment.
Identify the highest-leverage technical bets
Evaluate the technical landscape and identify the most important investments that will unlock customer success, accelerate delivery, and create long-term leverage for the vertical. Prioritize decisions thoughtfully, balancing immediate customer needs against durable architecture.
Ship the first version of the vertical's infrastructure
Build the initial technical foundation for the vertical by leveraging the core platform where it fits and extending it where the vertical requires new capabilities. Deliver production-grade systems, integrations, workflows, and tooling that create a foundation future engagements can build upon.
Lead customer engagements end-to-end
Own early customer engagements from technical scoping through delivery and post-launch support. Use learnings from those engagements to continuously improve the underlying infrastructure, implementation patterns, and technical playbook.
Create reuse and drive productization
Successfully reuse the vertical's technical foundation across multiple customer engagements, demonstrating that the systems being built are durable rather than one-off solutions. Surface repeatable patterns, infrastructure gaps, and product opportunities, and partner with Product and Engineering to elevate proven capabilities into the core platform.
What You Bring
Must Haves
3+ years of engineering experience, including meaningful 0→1 work as a founding engineer, early technical lead, or builder in a highly ambiguous environment.
Strong engineering generalist instincts with a backend and data orientation.
Hands-on experience with Python and SQL.
Comfort working across infrastructure, application logic, and data systems.
Ability to create structure where none exists and move quickly without a fully defined roadmap.
Strong technical judgment, especially around short-term delivery versus long-term architecture.
Strong written and verbal communication skills, including the ability to work directly with senior technical and business stakeholders.
Ability to independently run technical scoping conversations with customers and translate them into concrete execution plans.
Nice to Haves
Prior founding engineer experience at a successful startup.
Experience extending an existing platform into a new domain or use case.
Track record of turning customer-specific work into reusable internal infrastructure or product capabilities.
Familiarity with ML, NLP, or LLM-based systems.
Why This Role Is Special
This is a rare opportunity to define the technical foundation of a new business line inside a company that already has a real platform, real customers, and real momentum.
You won’t just be executing against a spec. You’ll help decide what the spec should be. If you enjoy building in ambiguity, working directly with customers, and creating systems that become the basis for an entire vertical, this role is for you.
Protege Values
Pass the Loved Ones’ Test
We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
Forward Deployed Engineer, Healthcare
The full posting opens here — pay, setting and the full description, without leaving the list.