fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role
This is a founding role on the small team that owns fal's core product systems — the connected tissue the whole business runs on: usage, billing, pricing, accounts, access, API keys, and model discovery. You'll join early, own these systems end-to-end (from data model through rollout to operational health), and design the APIs and product surfaces other teams and customers depend on. These are backend systems where correctness underpins real revenue and customer trust.
You will inherit systems that are functional, but you'll own their evolution over time to meet our growing customer needs and scale. Your first priority will be billing, fraud, and pricing. These are our highest-impact business initiatives, directly affecting both customer experience and revenue. From there, you'll expand your ownership across accounts, access and identity, API keys, and model discovery.
If you enjoy building and evolving high-impact platform systems that directly influence both our customers' experience and the effectiveness of our internal teams, this is the place to be.
What you'll do
Own billing and pricing first — metering, usage roll-ups, pricing logic, and the customer/admin surfaces that make cost understandable and trustworthy.
Design APIs, data models, and product contracts other teams safely build on.
Grow into the broader portfolio — accounts, access/identity, API keys, and model discovery — and the surfaces that expose usage, cost, and access.
Engineer for correctness and resilience: source of truth, idempotency, auditability, permissions, observability, graceful failure.
Own operational health, including on-call for the systems you build.
Partner across product, infra, finance, revenue, security, and customer-facing teams — turning messy, manual workflows into durable systems.
What we're looking for
5+ years building reliable, production backend systems (Python and/or TypeScript).
API & data-model design for complex product or business domains.
Can ship product and internal surfaces in React / Next.js / TypeScript.
Strong systems judgment (source of truth, idempotency, auditability, permissions, observability) and product judgment (workflows, edge cases, customer trust).
Thrives in ambiguous, cross-functional work where product, business logic, and infra overlap; ships a pragmatic v1 that can evolve. Priorities move fast — you're comfortable with that.
Identity & access — permissions, SSO, API keys, org/account hierarchy.
Developer platforms, model registries, or usage/observability dashboards.
Generative AI / inference APIs or other developer-facing AI products.
Our stack
Python (primary backend) · TypeScript + React/Next.js (product surfaces) · PostgreSQL / SQL · with some migration to Rust underway.
The team & who you'll work with
You'll be an early member of a small, high-leverage team (currently ~4) that owns fal's core product and billing systems. You'll work daily with product and infra engineers, and partner closely with finance, revenue, security, and customer-facing teams — this role lives at the seam of engineering and the business, so cross-functional partnership is core to the job, not a side task.
What fal offers
Interesting, challenging work at the frontier of generative media · competitive salary and equity · learning and growth · relocation assistance + visa sponsorship to San Francisco · health, dental, and vision · regular team events and offsites.
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role:
As a Senior Data Engineer at fal, you will build the data infrastructure that turns our internal systems and external vendor relationships into a clear picture of cost, margin, and performance. Your work spans both edges of our stack - the production infrastructure that runs every model invocation, and the partner APIs and compute vendors whose costs we need to reason about in near real-time.
This role sits at the intersection of software engineering and data engineering. You'll partner closely with Infra to safely instrument core systems, design a low-latency analytical write path, and stand up the ingestion pipelines that unlock cost, margin, and infrastructure analytics for the entire company. You will be a force multiplier - freeing up product engineers and infra to focus on what they do best while giving the data team the foundations it needs to move fast.
What you'll do
Instrument fal's core infrastructure to capture CPU, GPU, and request-level signals, working alongside our infra team to land changes safely in critical paths.
Build ingestion pipelines from partner APIs, compute vendors, and internal services into BigQuery and a new low-latency analytical store (e.g., ClickHouse).
Design and operate the ETL backbone that powers cost, margin, and usage analytics with durable, observable pipelines.
Stand up a lightweight, low-latency write path that the data team and product engineers can target directly for analytics-grade telemetry.
Partner with infra, data and product engineering to define data contracts and instrumentation standards, and act as the connective tissue between operational systems and the analytics layer.
What we are looking for
5+ years of experience as a software or data engineer, with a software-engineering-heavy track record (Python, Go, or similar)
Demonstrated ability to ship code into critical production infrastructure safely, including familiarity with database performance, query patterns, and incident risk.
Hands-on experience building ingestion pipelines into a warehouse (BigQuery, Snowflake, Redshift) and at least one low-latency analytical store (ClickHouse, Druid, Pinot, or similar).
Strong SQL and working proficiency in dbt and orchestration tooling (Dagster, Airflow, Prefect).
Track record of partnering across teams (infra, product engineering, data) and translating business questions into durable systems.
Bias for action and comfort working in fast-moving, ambiguous environments.
Nice-to-haves
Experience instrumenting GPU/accelerator workloads or other infrastructure-cost-heavy systems.
Exposure to FinOps or infrastructure cost modeling at a cloud-native company.
Experience with developer-facing API products or platforms.
Early-stage or fast-scaling startup experience.
Location
San Francisco, CA
What we offer at fal
Interesting and challenging work
A lot of learning and growth opportunities
We are currently hiring in downtown San Francisco.
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Staff Security Engineer, Infrastructure
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Engineering
Security
Compensation
$180K – $250K • Offers Equity
Overview
Application
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About the Role
We’re looking for a Security Engineer, Infrastructure to secure the core systems that power fal.ai’s platform: GPU compute, multi-cloud environments, networking, and data pipelines. You’ll operate across the full stack, from cloud and Kubernetes to identity, networking, and secrets, designing and implementing security controls that scale with a high-performance AI platform. This role is highly hands-on and systems-oriented, sitting at the intersection of security, infrastructure, and distributed systems.
What You’ll Do
Build & Harden Infrastructure Security
Design and implement security controls across:
Cloud infrastructure
Kubernetes and containerized workloads
Networking, service meshes, and edge systems
CI/CD pipelines and deployment systems
Secure compute environments for GPU workloads and model execution
Identity, Secrets & Access
Machine identity and workload authentication
Secrets management and encryption (e.g., Vault, KMS)
Least-privilege access and short-lived credentials
Implement Zero Trust principles across infrastructure
Secure AI & Data Systems
Protect model weights, inference endpoints, and customer data
Design secure data access pathways and isolation mechanisms
Ensure safe multi-tenant execution environments
Automation & Security Tooling
Build security guardrails directly into infrastructure and CI/CD
Use Infrastructure-as-Code (Terraform, Pulumi) to enforce secure defaults
Continuously identify and remediate security gaps through automation
Threat Modeling & Risk Reduction
Identify and mitigate risks across infrastructure layers
Defend against both external attackers and insider threats
Drive projects like network isolation, encryption, and secure service communication
Cross-Functional Collaboration
Partner with platform, infra, and ML teams to drive shift-left security
Enable engineers to move fast with secure-by-default systems
Contribute to a strong security culture across the company
What We’re Looking For
Core Requirements
8+ years in security engineering, infrastructure, or SRE
Strong understanding of:
Cloud security (AWS, GCP, or Azure)
Networking fundamentals (segmentation, firewalls, Zero Trust)
Linux systems and container security (Docker, Kubernetes)
Experience building or securing production infrastructure at scale
Security Expertise
Deep knowledge of:
Authentication & authorization systems
Secrets management and cryptography basics
Common vulnerabilities and attack vectors
Ability to design security controls across multiple layers (infra → app)
Engineering Skills
Proficiency in at least one language (Go, Python, or similar)
Experience with Infrastructure-as-Code (Terraform preferred)
Strong automation mindset—security should scale with systems
Nice to Have
Experience with:
GPU infrastructure or ML systems
Multi-tenant platform isolation
Service mesh / zero-trust architectures
High-growth startup environments
What Makes This Role Unique
Work on cutting-edge AI infrastructure security (not just SaaS)
Secure GPU clusters, model execution, and real-time inference systems
High ownership: design systems from first principles
Direct impact on developer trust and platform reliability
Our Security Philosophy
Secure-by-default > bolt-on security
Enable developers, don’t block them
Automate everything
Assume breach, design for resilience
Compensation & Benefits
Competitive salary + equity
Full health, dental, and vision coverage
Opportunity to work on frontier AI infrastructure
Why fal.ai
You’ll help define what security looks like for the next generation of AI infrastructure—where performance, scale, and safety all matter.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Software Engineer, Platform
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Engineering
Platform
Compensation
$180K – $250K • Offers Equity
Overview
Application
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.
Key responsibilities
Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc
Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
Leverage AI to an extreme level to build tools and automate alerting and recovery
Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage
Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)
Develop a suite of automated error detection and recovery processes
Work with partners to solve technical issues
Requirements
3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
Strong software engineering skills in Python; you write production tooling, not scripts
Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
Experience building internal tools or dashboards for infrastructure visibility
Excellent communication and ability to drive technical decisions across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)
Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2
Experience with AMD GPUs
Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)
Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)
Compensation
$180,000-250,000 plus equity + benefits
Location
San Francisco, CA
What we offer at fal
Interesting and challenging work
A lot of learning and growth opportunities
We are offering relocation assistance to San Francisco.
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Software Engineer, Site Reliability
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Engineering
Compensation
$180K – $250K • Offers Equity
Overview
Application
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role
You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.
Key Responsibilities
Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
Build and maintain CI/CD pipelines and deployment infrastructure
Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
Build dashboards, alerting, and anomaly detection across our systems
Define and enforce SLOs and build out incident response processes
Manage and improve our networking, load balancing, and service mesh configurations
Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
Requirements
5+ years experience in managing critical production systems and software development workflows
Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
Proficiency in Python and either Go or Bash for tooling and automation
Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
Excellent communication and ability to drive technical decisions across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
Experience with managing GPU and AI/ML workloads
Experience with kernel-based monitoring and routing (eBPF, XDP)
Experience with security tooling (Falco, Coroot, SIEM)
Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
Experience with distributed storage systems (Ceph, Longhorn, etc.)
Location
San Francisco, CA
What we offer at fal
Interesting and challenging work
A lot of learning and growth opportunities
We are currently hiring in downtown San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Software Engineer, Full Stack (Serverless)
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Engineering
Compensation
$180K – $230K • Offers Equity
Overview
Application
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role:
As a Full Stack Engineer on Serverless, you will build the core product across frontend and backend that powers fal’s Serverless platform. This is a deeply product-focused role. You will work side-by-side with Product and Infrastructure to design and ship reusable, scalable systems that enterprise customers rely on in production every day.
You will be a foundational technical owner of fal Serverless as it scales to thousands of enterprise customers, with real responsibility, autonomy, and impact. This is a chance to help build a new product vertical from the ground up inside a company that is already scaling at rocket-ship speed.
What you’ll work on:
Build and maintain core Serverless UI features (dashboards, logs, observability, configuration, usage)
Design and implement backend APIs that power the Serverless product experience
Improve performance, reliability, and scalability of customer-facing systems
Work closely with Infrastructure to ensure product features align with platform capabilities
Own features end-to-end, from design through production and iteration
What we’re looking for:
Strong experience working across both frontend and backend
Proficiency with TypeScript, Python, Postgres, and Next.js
Experience owning features end-to-end in production systems
Ability to context switch between UI, backend, and performance work
Product-minded engineer who values clean abstractions and long-term maintainability
Comfortable working in a fast-moving, low-process environment
Nice to have:
Experience building developer platforms or infrastructure-adjacent products
Familiarity with observability tooling (logging, metrics, tracing) in production environments
Background in distributed systems, container orchestration, or cloud-native architectures
Experience with real-time systems, streaming logs, or high-throughput data pipelines
Exposure to technologies such as Kubernetes, Prometheus, Datadog, gRPC, or similar systems
Entrepreneurial mindset and strong ownership mentality
What we offer at fal:
Interesting and challenging work
Competitive salary and equity
A lot of learning and growth opportunities
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsite
Location:
San Francisco, CA
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Senior Software Engineer, Product
Location
San Francisco
Employment Type
Full time
Location Type
On-site
Department
Engineering
Compensation
$180K – $250K • Offers Equity
Overview
Application
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role
You are a versatile engineer who thrives on building and deploying seamless user experiences. You possess a strong understanding of both backend and frontend technologies, enabling you to take ownership of features from concept to launch. You are proficient in crafting robust APIs, managing databases, and developing interactive user interfaces. Your focus is on delivering high-quality, scalable, and maintainable products.
Key Responsibilities:
You will have access to our cloud infrastructure for development and deployment. You will make our model playgrounds more interactive and help make them more discoverable.
Some core technologies we use include Typescript, Python, Postgres, and Next.js.
You'll collaborate with a cross-functional team to rapidly iterate and deploy new features.
Requirements:
5+ years of full-stack engineering experience delivering reliable, scalable products in production
Extensive experience writing maintainable, production-quality code in TypeScript, Python, and JavaScript
Deep understanding of relational databases, including PostgreSQL, schema design, and performance optimization
Skilled in owning feature development from concept to launch, with an emphasis on reusable design systems and UI components
What we offer at fal:
Interesting and challenging work
Competitive salary and equity
A lot of learning and growth opportunities
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsite
Location:
We are currently hiring in downtown San Francisco.
Apply for this Job
Powered by
Privacy PolicySecurityVulnerability Disclosure
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role:
You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data. You know how to achieve reliability and scale with minimum operational load.
Key responsibilities
Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc
Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world
Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems
Profile and tune low level CPU and memory performance
Requirements
3+ years experience building distributed compute and orchestration platforms in Python or Rust
Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning
Deep understanding of computational complexity and memory allocation
Track record of designing systems that scale under real production load
Experience building and using observability to drive performance and reliability decisions
Excellent communication and ability to drive technical decisions across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
Experience with AI/ML inference or training infrastructure
Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)
Background in building multi-tenant compute platforms
Understanding of networking fundamentals and performance characteristics
Familiarity with GPU workload characteristics and scheduling constraints
Compensation
$180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)
Location
San Francisco, CA
What we offer at fal
Interesting and challenging work
A lot of learning and growth opportunities
We are currently hiring in downtown San Francisco.
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites
Listed by Fal for a position based in the United States. Employers on this board attest they are hiring domestically.
Select a role
The full posting opens here — pay, setting and the full description, without leaving the list.