Our mission is to architect AI that learns from and interacts with the world like humans do.
We're pioneering the model architectures that will make this possible. Our founding team met as PhDs at the Stanford AI Lab, where we invented State Space Models or SSMs, a new primitive for training efficient, large-scale foundation models. Our team combines deep expertise in model innovation and systems engineering paired with a design-minded product engineering team to build and ship cutting edge models and experiences.
We're funded by leading investors at Index Ventures and Lightspeed Venture Partners, along with Factory, Conviction, A Star, General Catalyst, SV Angel, Databricks and others. We're fortunate to have the support of many amazing advisors, and 90+ angels across many industries, including the world's foremost experts in AI.
The Role
We are looking for a Product and Research Operations Manager to design, scale, and operate Cartesia's global scaled evaluation workforce. This role sits at the intersection of product operations, data operations, and vendor management, and directly impacts model quality and customer outcomes. You will own the end-to-end workforce system: hiring pipelines, vendor strategy, workforce planning, quality control, and operational performance.
You are building a production system of humans-in-the-loop for AI, translating ambiguous product needs into operational workflows and partnering across product, engineering, data, and customer-facing teams to support real-world evaluation at scale.
Your Impact
Design and implement workforce structure across languages, skill tiers, and use cases, including evaluators, auditors, and leads for TTS products
Build capacity models to support continuous eval pipelines and data production workflows
Own relationships with vendors such as data annotation firms and contractor platforms, negotiating rate cards, SLAs, and throughput guarantees
Decide on build, buy, or hybrid workforce models and continuously benchmark cost and performance across regions
Design multi-layer QA systems spanning self-checks, peer review, audits, and gold tasks
Define and track inter-rater reliability, error rates by category, and annotator-level performance distributions
Build escalation and retraining workflows to maintain quality at scale
Run day-to-day operations including task allocation, throughput tracking, and SLA adherence
Build systems to reduce evaluator fatigue, rotate task types, and maintain consistency across large-scale evaluations
Partner with tooling teams to improve evaluator UX and with data teams to ensure clean, structured outputs for model training
What You Bring
6+ years in operations, workforce management, or data annotation systems
Experience managing large contractor or vendor-based workforces
Proven ability to scale operations from zero to production
Systems thinking with the ability to design scalable operational frameworks
Strong analytical skills with comfort around metrics like inter-rater reliability, precision, and throughput
Ability to execute quickly under ambiguity with close attention to quality and edge cases
Nice to Have
Experience in AI/ML data operations or evaluation pipelines
Background in audio, speech, or language-related workflows
Familiarity with QA systems and annotation tooling
Experience with marketplace platforms such as Upwork or Mercor
Exposure to multilingual operations
Note: Cartesia participates in E-Verify and will provide the federal government with Form I-9 information to confirm employment eligibility after hire.
More Details
🏢 In-office policy: We’re an in-person team based out of offices in 🇺🇸 San Francisco, 🇬🇧 London and 🇮🇳 Bangalore. We love being in the office, hanging out together, and learning from each other every day.
🌎 Visa sponsorship: We provide visa sponsorship support and assess each circumstance on a case-by-case basis. However, visa sponsorship is dependent on many factors, including the role you are applying for, and the location you are going to be based, and so we can't always guarantee success. Your Recruiter will work with you to understand your visa sponsorship needs from the first call.
🚢 We ship fast. All of our work is novel and cutting edge, and execution speed is paramount. We have a high bar, and we don’t sacrifice quality or design along the way.
🤝 We support each other. We have an open & inclusive culture that’s focused on giving everyone the resources they need to succeed.
Our Benefits (US Employees Only)
💰 Compensation Competitive base salary alongside attractive equity package.
🩺 Health Insurance Fully covered medical insurance along with dental and vision for you and your family.
🚆 Commuter Allowance A monthly stipend to help you get to and from the office.
🏖️ Flexible PTO Take as much time as you need to recharge your batteries.
🍲 Meals & Snacks Lunch, dinner and plenty of snacks, provided daily.
🦖 Your own personal Yoshi
Our Commitment to Equal Opportunity
Cartesia is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, national origin, age, disability, veteran status, genetic information, or any other legally protected status.
Listed by Cartesia for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
91 days ago
Product & Research Operations Manager, Human Data
Cartesia · San Francisco, California, United States
ABOUT CARTESIA
Our mission is to architect AI that learns from and interacts with the world like humans do.
We're pioneering the model architectures that will make this possible. Our founding team met as PhDs at the Stanford AI Lab, where we invented State Space Models or SSMs, a new primitive for training efficient, large-scale foundation models. Our team combines deep expertise in model innovation and systems engineering paired with a design-minded product engineering team to build and ship cutting edge models and experiences.
We're funded by leading investors at Index Ventures and Lightspeed Venture Partners, along with Factory, Conviction, A Star, General Catalyst, SV Angel, Databricks and others. We're fortunate to have the support of many amazing advisors, and 90+ angels across many industries, including the world's foremost experts in AI.
The Role
We are looking for a Product and Research Operations Manager to design, scale, and operate Cartesia's global scaled evaluation workforce. This role sits at the intersection of product operations, data operations, and vendor management, and directly impacts model quality and customer outcomes. You will own the end-to-end workforce system: hiring pipelines, vendor strategy, workforce planning, quality control, and operational performance.
You are building a production system of humans-in-the-loop for AI, translating ambiguous product needs into operational workflows and partnering across product, engineering, data, and customer-facing teams to support real-world evaluation at scale.
Your Impact
- Design and implement workforce structure across languages, skill tiers, and use cases, including evaluators, auditors, and leads for TTS products
- Build capacity models to support continuous eval pipelines and data production workflows
- Own relationships with vendors such as data annotation firms and contractor platforms, negotiating rate cards, SLAs, and throughput guarantees
- Decide on build, buy, or hybrid workforce models and continuously benchmark cost and performance across regions
- Design multi-layer QA systems spanning self-checks, peer review, audits, and gold tasks
- Define and track inter-rater reliability, error rates by category, and annotator-level performance distributions
- Build escalation and retraining workflows to maintain quality at scale
- Run day-to-day operations including task allocation, throughput tracking, and SLA adherence
- Build systems to reduce evaluator fatigue, rotate task types, and maintain consistency across large-scale evaluations
- Partner with tooling teams to improve evaluator UX and with data teams to ensure clean, structured outputs for model training
What You Bring
- 6+ years in operations, workforce management, or data annotation systems
- Experience managing large contractor or vendor-based workforces
- Proven ability to scale operations from zero to production
- Systems thinking with the ability to design scalable operational frameworks
- Strong analytical skills with comfort around metrics like inter-rater reliability, precision, and throughput
- Ability to execute quickly under ambiguity with close attention to quality and edge cases
Nice to Have
- Experience in AI/ML data operations or evaluation pipelines
- Background in audio, speech, or language-related workflows
- Familiarity with QA systems and annotation tooling
- Experience with marketplace platforms such as Upwork or Mercor
- Exposure to multilingual operations
Note: Cartesia participates in E-Verify and will provide the federal government with Form I-9 information to confirm employment eligibility after hire.
MORE DETAILS
🏢 In-office policy: We’re an in-person team based out of offices in 🇺🇸 San Francisco, 🇬🇧 London and 🇮🇳 Bangalore. We love being in the office, hanging out together, and learning from each other every day.
🌎 Visa sponsorship: We provide visa sponsorship support and assess each circumstance on a case-by-case basis. However, visa sponsorship is dependent on many factors, including the role you are applying for, and the location you are going to be based, and so we can't always guarantee success. Your Recruiter will work with you to understand your visa sponsorship needs from the first call.
🚢 We ship fast. All of our work is novel and cutting edge, and execution speed is paramount. We have a high bar, and we don’t sacrifice quality or design along the way.
🤝 We support each other. We have an open & inclusive culture that’s focused on giving everyone the resources they need to succeed.
OUR BENEFITS (US EMPLOYEES ONLY)
💰 Compensation Competitive base salary alongside attractive equity package.
🩺 Health Insurance Fully covered medical insurance along with dental and vision for you and your family.
🧑🧑🧒🧒 Parental Leave 9 weeks paternity & 12 weeks maternity leave
🏦 401(k)
🚆 Commuter Allowance A monthly stipend to help you get to and from the office.
🏖️ Flexible PTO Take as much time as you need to recharge your batteries.
🍲 Meals & Snacks Lunch, dinner and plenty of snacks, provided daily.
🦖 Your own personal Yoshi
OUR COMMITMENT TO EQUAL OPPORTUNITY
Cartesia is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, national origin, age, disability, veteran status, genetic information, or any other legally protected status.
Listed by Cartesia for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
110 days ago
Principal AI Research Scientist, Research Director - AI Scaling
Principal AI Research Scientist, Research Director - AI Scaling, Mountain View, California; San Francisco, California. Join us! Together we can use data to solve the challenges of tomorrow
Listed by Databricks for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
110 days ago
Principal AI Research Scientist, Research Director - AI Scaling
Principal AI Research Scientist, Research Director - AI Scaling, Mountain View, California; San Francisco, California. Join us! Together we can use data to solve the challenges of tomorrow
Listed by Databricks for a position based in the United States. Employers on this board attest they are hiring domestically.
Science
131 days ago
Research Engineer, Multimodal Data
Eventual · San Francisco, California, United States
ABOUT EVENTUAL
Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. As a result, robotics and video-AI teams iterate on model improvement about once a week. Most of that week isn't training — it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year.
Eventual was founded in 2022 to close it. Our open-source engine, Daft https://daft.ai/, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop. Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate. One iteration per day becomes the norm.
We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc http://Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next.
Join our small (but powerful!) team working together 4 days/week in our SF Mission district office.
YOUR ROLE
As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content. Physical AI teams have raw footage, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. We change that economics: we run vision-language models over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes.
You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a research engineering role — meaning you'll read papers and run experiments, but you ship to production and your work is judged by what it does for customer training runs.
KEY RESPONSIBILITIES
- Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale.
- Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and convolutional perception models against customer datasets and benchmarks.
- Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical.
- Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs.
- Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering.
- Work directly with researchers at our partner labs — your shortest feedback loop is their next training iteration.
WHAT WE LOOK FOR
- Strong familiarity with modern vision and multimodal models — convolution nets, VLMs, VQA, embeddings — and a sense for the SOTA that's actually deployable today vs. on a leaderboard.
- Experience running these models at scale on real video and sensor data, ideally for perception tasks (detection, tracking, segmentation, retrieval, captioning).
- Background from a perception team at a self-driving, robotics, or visual-data company — or equivalent depth from a research lab.
- Comfortable with cloud infrastructure and large-scale data processing — you don't need to be a distributed-systems engineer, but you've shipped jobs that ran on thousands of GPU-hours of video.
- Bias toward data and infrastructure: you reach for "annotate the whole corpus" before "fine-tune another model."
NICE TO HAVE
- Experience training vision or multimodal models from scratch (not just calling APIs).
- ML/AI research background — papers, citations, or a research org on your resume.
- Hands-on time with big-data frameworks like Spark, Ray, or Daft.
- Worked on embeddings, retrieval, or content-aware search at scale.
- Experience designing labeling taxonomies or running annotation programs.
PERKS & BENEFITS
- In-person, tight-knit team — 4 days/week in our SF Mission office.
- Competitive comp and meaningful startup equity.
- Catered lunches and dinners for SF employees.
- Commuter benefit.
- Team-building events and poker nights.
- Health, vision, and dental coverage.
- Flexible PTO.
- Latest Apple equipment.
- 401(k) plan with match.
If you're excited about being on the team that turns petabytes of raw video into the training data for the next generation of Physical AI, we'd love to talk.
Listed by Eventual for a position based in the United States. Employers on this board attest they are hiring domestically.