Machine Learning Engineer (Inference)

LB147
  • $300,000 - $350,000
  • San Francisco Bay Area
  • Permanent

ML Inference Engineer


Build the infrastructure that makes cutting-edge AI fast enough to work at scale.


We’re hiring an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that has already grown to 8-figure revenue.


The team is rebuilding its LLM inference stack from the ground up, solving challenging systems problems around GPU performance, distributed compute and real-time model serving.

This role is for engineers who enjoy going deep on performance, infrastructure, and optimisation.


The role

  • Build the infrastructure that serves large-scale LLM workloads
  • Push the limits of latency, throughput and GPU efficiency
  • Design distributed inference across single and multi-GPU systems
  • Improve GPU scheduling, orchestration and resource utilisation
  • Profile and remove bottlenecks across compute, memory and networking
  • Scale production workloads across Kubernetes and GPU clusters
  • Make low-level architecture decisions where milliseconds matter

What we're looking for

  • Strong Python and/or C++
  • Experience with distributed systems or high-performance computing
  • Knowledge of LLM inference and model serving
  • Experience optimising GPU-heavy workloads
  • Exposure to CUDA, NCCL or Triton
  • Strong understanding of PyTorch and modern ML infrastructure
  • Experience with vLLM, TensorRT-LLM, SGLang or similar
  • Knowledge of techniques such as quantisation, batching, KV caching and parallelism

Why join?

  • Tackle genuinely difficult AI infrastructure problems
  • Work on systems operating at real production scale
  • Join a fast-growing company already at 8-figure revenue
  • Significant ownership over a new inference architecture
  • Work at the intersection of LLMs, GPUs and distributed systems
Anna Heneghan Senior ML Research & Engineering Recruiter

Apply for this role