Machine Learning Engineer (Inference)
LB147
Posted: 16/09/2026
- $300,000 - $350,000
- San Francisco Bay Area
- Permanent
ML Inference Engineer
Build the infrastructure that makes cutting-edge AI fast enough to work at scale.
We’re hiring an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that has already grown to 8-figure revenue.
The team is rebuilding its LLM inference stack from the ground up, solving challenging systems problems around GPU performance, distributed compute and real-time model serving.
This role is for engineers who enjoy going deep on performance, infrastructure, and optimisation.
The role
- Build the infrastructure that serves large-scale LLM workloads
- Push the limits of latency, throughput and GPU efficiency
- Design distributed inference across single and multi-GPU systems
- Improve GPU scheduling, orchestration and resource utilisation
- Profile and remove bottlenecks across compute, memory and networking
- Scale production workloads across Kubernetes and GPU clusters
- Make low-level architecture decisions where milliseconds matter
What we're looking for
- Strong Python and/or C++
- Experience with distributed systems or high-performance computing
- Knowledge of LLM inference and model serving
- Experience optimising GPU-heavy workloads
- Exposure to CUDA, NCCL or Triton
- Strong understanding of PyTorch and modern ML infrastructure
- Experience with vLLM, TensorRT-LLM, SGLang or similar
- Knowledge of techniques such as quantisation, batching, KV caching and parallelism
Why join?
- Tackle genuinely difficult AI infrastructure problems
- Work on systems operating at real production scale
- Join a fast-growing company already at 8-figure revenue
- Significant ownership over a new inference architecture
- Work at the intersection of LLMs, GPUs and distributed systems
Anna Heneghan
Senior ML Research & Engineering Recruiter