Inference Serving - vLLM

LB137
  • -
  • Santa Clara, CA | Onsite
  • Permanent

I'm hiring for a Machine Learning Engineer – LLM Inference Serving role on behalf of a stealth AI infrastructure startup 🚀


Must-have: deep hands-on experience with vLLM or SGLang 🔬


We’re looking for engineers who have:

  • Made public GitHub contributions to vLLM/SGLang
  • Refactored inference framework internals to improve performance
  • Worked on KV Cache lifecycle management
  • Used Ray, Dynamo, or similar inference orchestration platforms


📍 Santa Clara, CA | Onsite


Focus: LLM inference, distributed serving, throughput/latency optimisation, KV cache, scheduling


Please reach out with links to relevant GitHub contributions or inference framework work 📩

Zee Uddin Researcher

Apply for this role