Research Engineer

LB188
  • $260,000
  • San Francisco, CA
  • Permanent

Research Engineer (Voice / Speech AI)


📍 San Francisco, CA (Onsite)


I am seeking a Research Engineer to join a frontier AI research lab building the next generation of real-time conversational voice systems. Rather than building traditional AI assistants, this team is creating full-duplex AI companions that can naturally participate in live multiplayer experiences: listening, responding, interrupting appropriately, understanding shared context, and interacting alongside human players in real time.


Alongside building production voice agents, the team also operates as a voice research lab, producing research-grade speech datasets, evaluations, and tooling that support frontier AI research across the wider ecosystem.


What We're Looking For:


  • Experience researching speech foundation models or voice AI systems
  • Hands-on experience training, fine-tuning, evaluating, and benchmarking speech models
  • Experience working on streaming or low-latency speech systems
  • Exposure to full-duplex conversational models such as Moshi, PersonaPlex, or similar research is highly desirable
  • Publications at leading speech, audio, or machine learning conferences (ICASSP, Interspeech, NeurIPS, ICML, ICLR, etc.) are a significant advantage
  • Strong machine learning research background with the ability to translate research into production systems


What You'll Do:


  • Research and prototype ultra-low latency streaming and full-duplex speech systems
  • Develop speech foundation models, including cascaded (ASR → LLM → TTS) and speech-to-speech architectures
  • Improve conversational behaviours including turn-taking, interruption handling, timing, prosody, and emotion
  • Train, fine-tune, benchmark, and evaluate state-of-the-art speech models
  • Design evaluation frameworks that measure conversational quality and social presence—not just transcription accuracy
  • Work across model research, inference optimisation, data generation, and production deployment
  • Collaborate closely with researchers building frontier multimodal and voice foundation models


This is a genuinely research-focused opportunity for someone who wants to push the boundaries of speech AI. You'll be tackling challenges around conversational timing, streaming inference, interruption handling, latency optimisation, and voice quality—building systems that feel socially present rather than simply responsive.


Based in downtown San Francisco, the team works fully onsite and offers highly competitive compensation alongside meaningful equity as they continue to build one of the most ambitious voice AI research groups in the industry.


If you'd like to find out more, apply today!


Ethan Lewis

elewis@acceler8talent.com

Ethan Lewis Researcher

Apply for this role