Inference Search Partners · Confidential
Member of Technical Staff —
Model Optimization & Inference (New Grad)
Nuance Labs — Seattle, WA
Face-to-face AI that feels human — built from the ground up.
Most conversational AI avatars today are hacks — a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time. Nuance Labs is building something fundamentally different: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
Current systems take 2–5 seconds to respond. Natural conversation requires sub-500ms. That's a 10× improvement, and it demands rethinking the entire stack — developing foundation models designed for full-duplex from the ground up.
Research pedigree. Production ambition. No busy work.
We're a small, fast-moving research team with an exceedingly high bar — bringing on only the very best talent. Every member has massive ownership, deep trust, and the opportunity to shape both the technology and the company from the ground up.
Own inference, end to end.
You'll join a small team solving one of the hardest latency problems in AI today. The system has to listen and speak simultaneously, perceive emotion in real time, and respond with a face that actually reflects it. You're responsible for making it fast enough to feel human.
We're looking for an early-career engineer (BS, MS, or PhD — or nearing graduation) who's excited about taking trained models and squeezing every last millisecond out of them. We don't require a PhD — we care about systems intuition, engineering chops, and the appetite to go deep.
Squeeze every last millisecond out of the stack.
- Optimize inference end-to-end across the full model stack — LLMs, audio models, and diffusion-based components
- Implement and tune KV cache strategies for long-context conversations: eviction policies, compression, memory-efficient attention
- Work with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and extend them for Nuance's specific workloads
- Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
- Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
- Apply quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without degrading quality
- Build internal tooling: profiling viewers, end-to-end inference test harnesses, and infrastructure that helps the team move quickly
- Work closely with research and infrastructure to ensure new models ship with optimized serving from day one
Systems intuition. Engineering depth. Appetite to go deep.
Nuance isn't looking for someone who's only fine-tuned models. They want someone who understands — or wants to deeply understand — the full stack from model weights to serving infrastructure.
- 0–2 years of experience as an ML engineer (new grad level)
- Hands-on experience profiling and optimizing LLM or diffusion models (NSight, torch profiler)
- Experience productionizing ML models for inference and deployment at scale
- Familiarity with vLLM, SGLang, or similar inference serving frameworks
- Prior experience at a VC-backed startup or an AI team at top-tier big tech (Meta FAIR, Google DeepMind, Databricks) or strong LLM inference / ML systems research
- BS, MS, or PhD in CS, ML, or a related field — completed or nearly done
- Willing and able to work on-site in Seattle, 5 days a week
Small team. High bar. Full ownership.
Integrity and respect aren't negotiable.
Open communication, shared context — everyone can make informed decisions.
Bias toward action. Iterate fast. Learn quickly.
The best ideas happen face-to-face. Five days a week in Seattle.
Fast, technical, and respectful of your time.
Interested? Let's talk.
15 minutes is all it takes to find out if this is worth exploring further.
Book a Call with Shwetha