ABOUT ARENA INTELLIGENCE
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
THE ROLE
This is a senior individual contributor role for a closer who thrives on autonomy. You will own the full lifecycle of revenue generation—from navigating complex technical stakeholders to structuring and closing seven-figure contracts.
You will be responsible for ensuring Arena meets and exceeds its annual revenue and consumption targets by deeply embedding our evaluation infrastructure into the workflows of the world's most advanced AI companies.
CORE RESPONSIBILITIES
STRATEGIC ACCOUNT GROWTH & EXECUTION
- Own the Model Lab Relationships: Serve as the primary business lead for top global model labs, managing the relationship from research lead to procurement.
- Drive Revenue Expansion: Expand revenue across modalities (text, code, multimodal, vision, auto-eval) by aggressively aligning our eval programs with each lab’s training and research roadmaps.
- Strategic Enterprise Expansion: Identify and close select opportunities to bring LMArena’s eval capabilities to high-value enterprise accounts (e.g., Salesforce, IBM, ServiceNow) seeking to evaluate internal or vendor models.
- Secure Market Presence: Negotiate co-release and marketing rights for every major lab partnership to ensure LMArena retains mindshare as the industry standard.
- Lifecycle Management: Own contract renewals, upsells, and expansions, ensuring forward visibility into multi-quarter revenue commitments.
FORECASTING, DEAL STRUCTURING & STRATEGY
- Pipeline Hygiene: Maintain rigorous revenue forecasts that tie directly to delivery schedules, consumption pacing, and renewal cycles.
- Deal Architecture: Design and enforce pricing frameworks that balance predictability (committed contracts) with flexibility (usage-based consumption).
- Cross-Functional Alignment: Collaborate with Delivery and Product teams to ensure evaluation capacity, quality, and timing are aligned with customer expectations.
- New Product GTM: Ramp up on net new Eval or Data Products and build the strategy to sell them into existing and net new relationships.
WHAT WE LOOK FOR
- 10+ Years of Business Development / Sales: You have a proven track record of selling complex technical infrastructure or data products as a top-performing individual contributor.
- Proven Closer: You have a history of personally closing seven-figure deals. You are comfortable navigating procurement at large organizations and negotiating complex terms without hand-holding.
- AI Fluency: You have ideally worked in AI and understand the stack (Training vs. Inference, RLHF, Benchmarking). You can credible discuss model performance needs with research scientists.
- High Agency: You are a self-starter who can operate without a playbook. You effectively manage your own pipeline, forecasting, and deal hygiene.
- Operational Rigor: You understand how to align sales promises with delivery reality, working closely with finance and product to ensure deals are profitable and executable
WHAT WE OFFER
- We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
- Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
- The opportunity to work on cutting-edge AI with a small, mission-driven team
- A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
ABOUT ARENA INTELLIGENCE
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence builds the evaluation layer for AI, and our customers are the labs and enterprises building at the frontier. We're seeking a GTM Strategy and Operations Lead to help shape where we grow and to own the commercial foundation that supports it.
You'll work with the founders, go-to-market leadership, and finance on the decisions that set our trajectory: how we run pipeline and forecast, how we support the sales team, how we price and package, which markets to pursue. You'll also own the infrastructure underneath them, from the reporting that sales and finance both work from through deal structuring and quote-to-cash.
Reporting to the Head of Finance, you'll join a small team where most of this is still being built. We're looking for someone comfortable moving between a strategy recommendation and the data model behind what an account actually consumed, and who finds that combination interesting rather than daunting.
You'll
- Run the go-to-market operating cadence, including pipeline reviews, forecast discipline, and the business reviews that give sales leadership a clear read on performance.
- Define and build the pipeline and usage reporting that GTM and finance both work from.
- Be the operational partner to the sales team, owning the analysis, account-level visibility, and process support they need.
- Set pricing and packaging strategy and carry it through to execution, covering rate card design, discount guardrails, and how new products get monetized before they ship.
- Develop the go-to-market approach for new segments, from sizing the opportunity and building the investment case through the coverage model, packaging, and process needed to sell into them.
- Run the deal desk, structuring and pressure-testing large commercial agreements alongside the go-to-market team, and turning what gets signed into repeatable process.
- Own quote-to-cash from order form through metering, consumption tracking, and billing accuracy, in partnership with accounting.
- Partner with product and engineering so new SKUs launch instrumented, billable, and correctly represented in systems.
You'll have
- 6+ years of experience across GTM strategy, revenue operations, sales strategy, deal desk, or management consulting, including time at a high-growth technology company selling to enterprises.
- Experience running go-to-market operating rhythms (pipeline reviews, forecasting, and performance reporting) as the operational partner a sales team relies on.
- Direct experience with pricing and packaging decisions, and fluency in consumption or usage-based revenue models.
- A track record of structuring ambiguous commercial problems, including large non-standard agreements, and staying with them through implementation into systems and process.
- Working knowledge of quote-to-cash mechanics, including where the handoffs between CRM, billing, and accounting usually break.
- Comfort pulling your own data using SQL or a BI tool.
- Fluent with AI and vibe coding in your daily workflow, with the judgment to know when the answer requires doing the work yourself.
- Strong written and verbal communication, with the ability to hold a position with sales leadership and executives.
- High ownership and sound instincts about when a process is worth building.
WHAT WE OFFER
- We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
- Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
- The opportunity to work on cutting-edge AI with a small, mission-driven team
- A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
What You'll Do
Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
You’ll have
4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome and Orb
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
What You'll Do
Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
You’ll have
4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome and Orb
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
ABOUT ARENA INTELLIGENCE
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
ABOUT THE ROLE
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
WHAT YOU'LL DO
- Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
- Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
- Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
- Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
YOU’LL HAVE
- 4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
- Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
- Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
- A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
- Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
NICE TO HAVE
- Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
- Background in ML infrastructure, model serving, or evaluation frameworks.
- Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
- Experience building billing infrastructure around systems like Stripe, Metronome and Orb
- Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
WHAT WE OFFER
- We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
- Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
- The opportunity to work on cutting-edge AI with a small, mission-driven team
- A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Site Reliability Engineer
Location
Bay Area
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Overview
Application
About Arena Intelligence
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
What You'll Do
Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
What We're Looking For
6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome and Orb
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Backend Engineer
Location
Bay Area
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Overview
Application
About Arena Intelligence
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for a backend engineer to build the products and platforms that sit on top of our evaluation systems - the APIs, services, and data systems that turn Arena’s evals into products that enterprises depend on.
This is the application layer of the Arena Service. Where our infrastructure teams builds the gateways and runtimes beneath - you will work closely with them to build the backends that customers will touch directly. This will include the APIs that expose Arena and enterprise data and the systems that manage usage, billing and multi tenant access. This work spans product surfaces of varying maturity and the data pipelines behind them - zero to one in places and scale-it up in others.
You'll be an early member of our enterprise team, working closely with infrastructure engineers, researchers, and product leadership. We move fast and stay rigorous.
What You'll Do
Build API-based products from the ground up. Design and ship the low-latency, high-reliability APIs and services that power Leaderboards, Evals, and Arena data products — clean, versioned, and built for the developers who consume them.
Turn evaluations into product. Partner with the research team to take novel eval methods and make them durable, full-featured products: scoring pipelines, data models, and the APIs that expose them.
Ship enterprise-grade backends. Usage metering, cost attribution, billing integration, authentication, RBAC, multi-tenancy, and audit logging — the systems enterprise customers expect.
Own the data architecture. Unify our public and private evaluation data, design schemas that hold up as the product grows, and make Arena’s data queryable, consistent, and fast.
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
What We're Looking For
5+ years of backend engineering experience, with meaningful time spent on building product facing APIs, services and data systems at scale
Strong proficiency in a modern backend language — Go preferred — and the judgment to design APIs other engineers and customers will live with for years.
Solid data fundamentals. You’re comfortable modeling, querying, and scaling Postgres, and you know when to reach for Redis, a queue, or a warehouse.
Experience composing services into products — integrating payments, auth, analytics, or data pipelines into something coherent and reliable.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask “why” before “how.”
Comfort with ambiguity. We’re a startup. Scope is fluid, context shifts, and you’ll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and the realities of building on them: streaming, token accounting, rate limits, model-specific quirks.
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing and usage infrastructure around systems like Stripe, Metronome, and Orb.
Familiarity with the modern AI stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
What You'll Do
Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
What We're Looking For
6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome and Orb
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Backend Engineer
Location
Bay Area
Employment Type
Full time
Location Type
Hybrid
Department
Engineering
Overview
Application
About Arena Intelligence
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
Arena Intelligence is looking for a backend engineer to build the products and platforms that sit on top of our evaluation systems - the APIs, services, and data systems that turn Arena’s evals into products that enterprises depend on.
This is the application layer of the Arena Service. Where our infrastructure teams builds the gateways and runtimes beneath - you will work closely with them to build the backends that customers will touch directly. This will include the APIs that expose Arena and enterprise data and the systems that manage usage, billing and multi tenant access. This work spans product surfaces of varying maturity and the data pipelines behind them - zero to one in places and scale-it up in others.
You'll be an early member of our enterprise team, working closely with infrastructure engineers, researchers, and product leadership. We move fast and stay rigorous.
What You'll Do
Build API-based products from the ground up. Design and ship the low-latency, high-reliability APIs and services that power Leaderboards, Evals, and Arena data products — clean, versioned, and built for the developers who consume them.
Turn evaluations into product. Partner with the research team to take novel eval methods and make them durable, full-featured products: scoring pipelines, data models, and the APIs that expose them.
Ship enterprise-grade backends. Usage metering, cost attribution, billing integration, authentication, RBAC, multi-tenancy, and audit logging — the systems enterprise customers expect.
Own the data architecture. Unify our public and private evaluation data, design schemas that hold up as the product grows, and make Arena’s data queryable, consistent, and fast.
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
What We're Looking For
5+ years of backend engineering experience, with meaningful time spent on building product facing APIs, services and data systems at scale
Strong proficiency in a modern backend language — Go preferred — and the judgment to design APIs other engineers and customers will live with for years.
Solid data fundamentals. You’re comfortable modeling, querying, and scaling Postgres, and you know when to reach for Redis, a queue, or a warehouse.
Experience composing services into products — integrating payments, auth, analytics, or data pipelines into something coherent and reliable.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask “why” before “how.”
Comfort with ambiguity. We’re a startup. Scope is fluid, context shifts, and you’ll wear many hats. That should sound exciting, not stressful.
Nice to Have
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and the realities of building on them: streaming, token accounting, rate limits, model-specific quirks.
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing and usage infrastructure around systems like Stripe, Metronome, and Orb.
Familiarity with the modern AI stack (vLLM, LiteLLM, LangChain, etc.).
What we offer
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
ABOUT ARENA INTELLIGENCE
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
ABOUT THE ROLE
Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
WHAT YOU'LL DO
- Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
- Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
WHAT WE'RE LOOKING FOR
- 6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
- Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
- Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
- A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
- Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
NICE TO HAVE
- Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
- Background in ML infrastructure, model serving, or evaluation frameworks.
- Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
- Experience building billing infrastructure around systems like Stripe, Metronome and Orb
- Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
WHAT WE OFFER
- We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
- Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
- The opportunity to work on cutting-edge AI with a small, mission-driven team
- A culture that values transparency, trust, and community impact
Come help build the space where anyone can explore and help shape the future of AI.
Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.
Listed by Arena for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.