Infrastructure Engineer

LB135
  • $150k to $390k base + equity
  • San Francisco, CA
  • Permanent

Member of Technical Staff, Infrastructure

On-site | San Francisco, CA | 5 days per week

$150k to $390k base + equity


I’m working with a well-funded AI infrastructure startup (Series A) building a cloud platform that runs inference workloads across GPUs, CPUs, and emerging accelerator architectures.

The team is solving a difficult infrastructure problem: making new different hardware accelerators usable through one reliable platform, without requiring customers to redesign their software stack for every accelerator.


This role will build the cluster infrastructure behind that platform. You’ll determine how new hardware is brought online, how compute fleets are provisioned and operated, and how production inference systems remain reliable as they scale.


You’ll work on problems such as:

  • Deploying production clusters across different accelerator architectures
  • Automating bare-metal provisioning, validation, upgrades, and fleet lifecycle management
  • Improving cluster scheduling, resource utilization, isolation, and capacity management
  • Building observability systems for faster debugging, incident response, and recovery
  • Making new accelerators production-ready across drivers, firmware, networking, and orchestration
  • Partnering with runtime, compiler, distributed-systems, networking, and hardware engineers.


Looking for engineers who have:

  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm
  • A track record of building observable, recoverable, and reliable production systems


This is an opportunity to join a small, highly technical team and build production infrastructure across multiple generations and types of AI hardware. The strongest candidates will be able to explain what they personally built, operated, measured, and debugged at scale.


Worth a confidential chat?

Anna Button Researcher

Apply for this role