Machine Learning Researcher, Audio
Machine Learning Researcher, Audio
Location: San Francisco, CA or Remote
About Bland
At Bland.com, our mission is to empower enterprises to build AI phone agents at scale. Based in San Francisco, we are a fast-growing team reimagining how customers interact with businesses through voice. We have raised $100 million from leading Silicon Valley investors, including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.
Voice is quickly becoming the primary interface between businesses and their customers. We are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.
The Role: Machine Learning Researcher, Audio
As a Machine Learning Researcher at Bland, you'll be working on foundational research and development across the core components of our voice stack: speech-to-text, large language models, neural audio codecs, and text-to-speech. Your work will define how our agents understand, reason, and speak in real time at enterprise scale.
This is not a narrow research role. You will take ideas from theory to large-scale training to production inference systems serving millions of calls per day. You will design new modeling approaches, validate them with rigorous experimentation, and collaborate with engineering teams to deploy them into real customer environments.
What You Will Do
Build and Scale Next-Generation TTS Systems
Design and train large scale text-to-speech models capable of expressive, controllable, human-sounding output.
Develop neural audio codec-based TTS architectures for efficient, high-fidelity generation.
Improve prosody modeling, question inflection, emotional expression, and multi-speaker robustness.
Optimize for real-time, low-latency inference in production.
Advance Speech-to-Text Modeling
Build and fine-tune large scale ASR systems robust to accents, noise, telephony artifacts, and code switching.
Leverage self-supervised pretraining and large-scale weak supervision.
Improve transcription accuracy for real-world enterprise scenarios, including structured extraction and conversational nuance.
Pioneer Neural Audio Codecs
Research and implement neural audio codecs that achieve extreme compression with minimal perceptual loss.
Explore discrete and continuous latent representations for scalable speech modeling.
Design codec architectures that enable downstream generative modeling and controllable synthesis.
Develop Scalable Training Pipelines
Curate and process massive audio datasets across languages, speakers, and environments.
Design staged training curricula and data filtering strategies.
Scale training across distributed GPU clusters focusing on cost, throughput, and reliability.
Run Rigorous Experiments
Design ablation studies that isolate the impact of architectural changes.
Measure improvements using both objective metrics and perceptual evaluations.
Validate ideas quickly through focused experiments that confirm or eliminate hypotheses.
What Makes You a Great Fit
Deep Research Foundations
Experience with self-supervised learning, multimodal modeling, or generative modeling.
Ability to derive new formulations and implement them efficiently.
Expertise in Voice Modeling
Hands-on experience building or scaling TTS, STT, or neural audio codec systems.
Familiarity with large scale speech datasets and real-world audio variability.
Strong intuition for audio quality, prosody, and conversational dynamics.
Systems and Hardware Awareness
Experience training and serving large models on modern accelerators.
Knowledge of inference optimization techniques, including quantization, kernel optimization, and memory efficiency.
Understanding of real-time constraints in telephony or streaming environments.
Experimental Rigor
Track record of designing controlled experiments and meaningful ablations.
Comfortable working with both offline benchmarks and live production metrics.
Ability to move quickly from hypothesis to validation.
Builder Mentality
Comfortable in fast-moving startup environments.
Strong ownership mindset from research through deployment.
Excited by ambiguous, unsolved problems.
How You Show Up
You treat unsolved problems as opportunities to invent new paradigms.
You identify the single experiment that can validate an idea in days, not months.
You measure everything and let data drive decisions.
You are obsessed with making voice agents sound truly human.
You use AI tools aggressively to amplify your own impact and accelerate research cycles.
Bonus Points
Experience with large scale distributed training.
Research publications or open source contributions in speech or language AI.
Background in real-time speech systems or telephony.
PhD in ML, AI, or a related field, or equivalent research impact.
Benefits and Compensation
Healthcare, dental, vision, all the good stuff
Meaningful equity in a fast-growing company
Every tool you need to succeed
Beautiful office in Levi's Plaza, SF with rooftop views
Competitive salary: $160,000 to $250,000
If you are energized by building and scaling TTS models, pioneering neural audio codecs, and pushing the boundaries of speech-to-text systems, we would love to hear from you.
Machine Learning Researcher, Multimodal LLMs
Machine Learning Researcher, Multimodal LLMs
Location: San Francisco, CA or Remote
About Bland
At Bland.com, our mission is to empower enterprises to build AI phone agents at scale. Voice is quickly becoming the primary interface between businesses and their customers, and we are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.
We’ve raised $100M from leading investors including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.
The Role
We are looking for someone to contribute to the development of our next-generation multimodal LLM stack, combining speech, text, tools, and real-time reasoning into a single unified system. You’ll be responsible for building industry-leading conversational AI models that power Bland's agent, and taking them all the way from idea to production.
At Bland, we're not just thinking about text modeling. You will define how our agents listen, think, and act in real time, integrating streaming audio, tool execution, and dynamic context into a single coherent system. You will take ideas from research through production systems serving millions of calls per day.
What Makes You a Great Fit
Strong LLM / Multimodal Background
Experience with LLMs, multimodal models, or speech-language systems
Deep understanding of prompting, fine-tuning, and alignment techniques
Familiarity with neural audio codecs and modern multimodal LLM techniques
Fast Experimental Loop
You can go from idea → dataset → experiment → conclusion in days
You know how to design experiments that actually answer the question
Product Intuition
Strong sense for what makes an interaction feel natural vs robotic
Ability to translate abstract modeling ideas into user-facing improvements
Builder Mentality
You take ownership from research through deployment
You thrive in ambiguous, fast-moving environments
You care about impact, not just elegance
How You Show Up
You think in systems, not just models
You obsess over latency, correctness, and real-world behavior
You are comfortable discarding ideas quickly when data disagrees
You push toward simple abstractions for complex problems
Bonus Points
Experience with real-time voice systems or conversational AI
Background in tool-using agents or agent frameworks
Experience with multimodal datasets (audio + text + actions)
Contributions to LLM or speech-related research or open source
Compensation & Benefits
Competitive salary: $180,000 – $260,000
Meaningful equity
Full healthcare, dental, vision
Office in Levi's Plaza, SF
High autonomy, high impact
Business Development Representative
About Bland
At Bland.com, our goal is to empower enterprises to make AI-phone agents at scale. Based out of San Francisco, we're a quickly growing team striving to change the way customers interact with businesses. We've raised $100 million from Silicon Valley's finest; including Emergence Capital, Scale Venture Partners, YC, the founders of Twilio, Affirm, ElevenLabs, and many more.
About the Role
This is a full-time role located in San Francisco, CA. As a Business Development Representative (BDR) at Bland, your primary responsibility will be driving pipeline and growing net-new business. You'll actively prospect, qualify, and pass meetings to our AEs to move deals forward.
We're a startup, and we need people who understand what that means. We aren't a super traditional team, but we are an extremely effective one. We love unique backgrounds, hard workers, and intelligent people who take pride in everything they do.
What Makes You a Great Fit
Preferably 1+ year of SaaS BDR experience
Startup experience is a plus, especially in Enterprise SaaS
Comfortable understanding and communicating around complex technical problems
Experience in business development and customer relationship management
High level of agency—able to figure things out, build new processes, and just get it done
Bachelor's degree is preferred but not required
-
Comfortable with cold outreach (calls, emails, LinkedIn, etc.)
Benefits and Pay:
Healthcare, dental, vision, all the good stuff
Meaningful equity in a fast-growing company
Every tool you need to succeed
-
Beautiful office in Levi's Plaza, SF with rooftop views
If you don't have the perfect experience, that's fine! We're a bunch of drop-outs and hackers. Working at a start-up is really hard. We work a lot and we figure things out on the fly.
OTE: $100,000
Senior Infrastructure Engineer
About Bland
At Bland.com, our goal is to empower enterprises to make AI-phone agents at scale. Based out of San Francisco, we're a quickly growing team striving to change the way customers interact with businesses. We've raised $100 million from Silicon Valley's finest; including Emergence Capital, Scale Venture Partners, YC, the founders of Twilio, Affirm, ElevenLabs, and many more.
About the Role
As a Senior Infrastructure Engineer at Bland, you'll help us to build the backbone that enables millions of AI-powered phone conversations. You're not just keeping servers running, you're architecting distributed systems that handle real-time voice processing, scale ML inference, and integrate with enterprise telephony infrastructure. Your work directly determines whether our platform can handle business-defining call volumes for our customers, or leaves them with dead air.
What You'll Do
Contribute to the designing of scalable architecture: Build distributed systems using Kubernetes that handle high-volume, real-time voice processing with strict latency and reliability requirements.
Build and Support ML infrastructure: Create and optimize the infrastructure supporting our AI models, from training pipelines to real-time inference serving across multiple regions.
Integrate with telephony: Maintain robust connections between our platform and complex enterprise phone systems, SIP trunks, and VoIP infrastructure.
Recognize Flaws, Control for them: We’re building a new type of architecture that takes something from Column A, and Column B. We’re never going to get it perfect, so you’ll be helping us keep a look out for what we need to solve.
Ensure reliability: Implement monitoring, alerting, and incident response systems that keep our platform running 24/7 with enterprise-grade uptime.
Scale with growth: Anticipate and solve scaling challenges before they become problems—our call volume grows exponentially and infrastructure needs to stay ahead.
-
Security and compliance: Implement security best practices and compliance requirements for enterprise customers in regulated industries.
Interesting Problems to Own
Old-Meets-New: Telephone calls have been around for awhile. Now, with an explosion in modern technologies, comes interesting new ways to wrangle old-school protocols and techniques. You’ll have the space to be creative and really own a new emergent type of architecture.
Sizable Call Volumes requires new approaches: Understand and deeply invest in ensuring that we match any amount of customer’s customers call volume! We need unique solutions, that you’ll help us discover along the way.
-
Streaming Architectures: On top of building to support our APIs, you’ll also be building to helping maintain the reliability, failover, and scaling of our important stream-based traffic.
What Makes You a Great Fit
Infrastructure expertise: 5+ years building and scaling distributed systems, with deep knowledge of cloud infrastructure (AWS/GCP preferred).
You “get” the fundamentals, and beyond: For example, you can casually tell someone how TLS works beyond buzzwords, do a quick sketch of how different load balancing strategies work, or even tell us the obscure thing you fell asleep reading about last night. There isn’t a blank stare, there’s an excitement to share.
Real-time systems experience: You've built systems that handle high-throughput, low-latency workloads, streaming, real-time processing, or similar.
Startup mentality: You've worked at fast-growing companies where you wear multiple hats and solve problems as they come up.
You’re opinionated, but you’re not alienating: You accept that opinions drive progress, but you don’t intend to break into alienating discussions at the risk of not finding compromises for our customers.
-
You’re familiar with some tools/components like: Cloudflare, HAProxy, Go, TypeScript, Datadog, Terraform, Docker, Kubernetes, Nvidia Hardware (nvlink for example), and anything in between.
Bonus Points If You Have
Experience with telephony systems (SIP, VOIP, WebRTC.)
Background in ML infrastructure, model serving, or GPU computing.
-
Experience with real-time audio/video processing.
Benefits and Pay:
Healthcare, dental, vision, all the good stuff
Meaningful equity in a fast-growing company
Every tool you need to succeed
-
Beautiful office in Levi's Plaza, SF with rooftop views
If you don't have the perfect experience that is fine! We're a bunch of drop-outs and hackers. Working at a start-up is really hard. We work a lot and we figure things out on the fly.
Compensation Range: $120,000-$200,000
The full posting opens here — pay, setting and the full description, without leaving the list.