Inference Engineer
- $200k base + equity
- San Francisco, CA
- Permanent
Member of Technical Staff, Inference Performance
On-site | San Francisco, CA | 5 days per week
200k base + equity
I’m working with a AI infrastructure startup building a high-performance inference cloud for open models.
The team optimizes the full path from model and serving engine through kernels, accelerators, and production infrastructure. Its platform is already processing trillions of tokens per month, and the company is expanding due to customer demand growing faster than its current capacity.
This role will focus on making inference faster, more reliable, and more cost-efficient across different models and hardware architectures.
You’ll work on problems such as:
- Profiling and optimizing inference latency, throughput, and memory usage
- Improving serving engines through batching, caching, quantization, and speculative decoding
- Developing and tuning CUDA, HIP, or Triton kernels
- Operating heterogeneous accelerator clusters
- Benchmarking models and hardware under realistic production workloads
- Building observability and reliability into the serving platform
- Debugging performance across models, runtimes, kernels, networking, and hardware.
Looking for engineers who have strong evidence in one or more of:
- High-performance AI inference or model-serving systems
- GPU kernel development using CUDA, HIP, or Triton
- Quantization, speculative decoding, batching, or KV-cache optimization
- PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks
- GPU, accelerator, HPC, or distributed-compute infrastructure
- Low-level performance profiling and systems optimization
- Operating latency-sensitive systems in production.
Strong candidates will be able to explain what they personally optimized, how they measured it, and the production impact it created.
This is an intense, highly hands-on environment with direct founder access and broad ownership. It will suit engineers who want to move across models, kernels, hardware, and infrastructure rather than remain within a narrowly defined area.
Worth a confidential chat?