Inference Search Partners · Exclusive Placement
Nuance Labs
Face-to-face AI interaction that feels human
Two ways to build here.
Both roles are on-site in Seattle, 5 days/week. Both come with relocation support and visa sponsorship. Both report directly into the founding team.
- PhD in ML / CS / RL completing or completed 2024–2026
- PPO, GRPO, DPO, RLHF — deep internals, not black box
- PhD research focused on frontier RL algorithms
- On-site Seattle, 5 days/week
- Internship at OpenAI, DeepMind, ByteDance Seed, or Mistral
- First-author at NeurIPS, ICML, or ICLR
- verl, OpenRLHF, SGLang, or vLLM experience
- Multimodal / audio / video RL work
- 2+ years ML engineering in production (not intern-only)
- Profiling with NSight or torch.profiler — hard dealbreaker
- Python + Rust or Go
- On-site Seattle, 5 days/week
- CUDA kernel optimization for inference acceleration
- Quantization: INT8, INT4, GPTQ, AWQ, FP8
- KV cache eviction, compression, memory-efficient attention
- Real-time audio/video systems (WebRTC, streaming)
Face-to-face AI that feels human — built from the ground up.
Most conversational AI avatars today are hacks — a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time. Nuance Labs is building something fundamentally different: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
Current systems take 2–5 seconds to respond. Natural conversation requires sub-500ms. That's a 10× improvement, and it demands rethinking the entire stack — developing foundation models designed for full-duplex from the ground up.
Research pedigree. Production ambition. No busy work.
A small, fast-moving research team with an exceedingly high bar — bringing on only the very best. Every member has massive ownership, deep trust, and the opportunity to shape both the technology and the company from the ground up.