Inference Search Partners · Confidential

Member of Technical Staff —
Model Optimization & Inference (New Grad)

Nuance Labs — Seattle, WA

$200K – $300K Series A · $60M On-site · Seattle, WA New Grad / Early Career Visa Transfers OK Relocation Covered

Face-to-face AI that feels human — built from the ground up.

Most conversational AI avatars today are hacks — a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time. Nuance Labs is building something fundamentally different: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.

Current systems take 2–5 seconds to respond. Natural conversation requires sub-500ms. That's a 10× improvement, and it demands rethinking the entire stack — developing foundation models designed for full-duplex from the ground up.

$60M Total Funding
25 People on Team
2024 Founded
<500ms Target Latency
Backed by Accel Lightspeed Ventures South Park Commons

Research pedigree. Production ambition. No busy work.

We're a small, fast-moving research team with an exceedingly high bar — bringing on only the very best talent. Every member has massive ownership, deep trust, and the opportunity to shape both the technology and the company from the ground up.

FM
Fangchang Ma
Co-founder & CEO
MIT PhD in Robotics & ML. Previously Research Manager at Apple, with experience at DJI. Published in top AI conferences.
2,400+ citations
EZ
Edward Zhang
Co-founder & CTO
UW PhD in Computer Graphics. Previously Senior Research Scientist at Apple, with experience at Google and Microsoft.
KY
Karren Yang
Founding Research Scientist
MIT PhD. Previously Senior Research Scientist at Apple, Niantic Labs, Meta Reality Labs, Bosch Center for AI, and Adobe Research.
1,800+ citations
CV
Claudia Vanea
Founding Research Scientist
Oxford PhD in AI. Previously founder at South Park Commons. Developed the first AI computational framework in her field — performance surpassing domain experts.
YS
Yaser Sheikh
Advisor
Former VP of Codec Avatars at Meta. Consulting Professor at Carnegie Mellon University's Robotics Institute — one of the world's foremost experts on photorealistic avatars.

Own inference, end to end.

You'll join a small team solving one of the hardest latency problems in AI today. The system has to listen and speak simultaneously, perceive emotion in real time, and respond with a face that actually reflects it. You're responsible for making it fast enough to feel human.

We're looking for an early-career engineer (BS, MS, or PhD — or nearing graduation) who's excited about taking trained models and squeezing every last millisecond out of them. We don't require a PhD — we care about systems intuition, engineering chops, and the appetite to go deep.

Base Salary
$200K – $300K
Depending on experience
Equity
Competitive
Series A — meaningful upside
Location
Seattle
On-site · 5 days/week
Relocation
Covered
Assistance available
Visa
Open
OPT & H1B transfers
Start
ASAP
Within 2 months preferred

Squeeze every last millisecond out of the stack.

  • Optimize inference end-to-end across the full model stack — LLMs, audio models, and diffusion-based components
  • Implement and tune KV cache strategies for long-context conversations: eviction policies, compression, memory-efficient attention
  • Work with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and extend them for Nuance's specific workloads
  • Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
  • Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
  • Apply quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without degrading quality
  • Build internal tooling: profiling viewers, end-to-end inference test harnesses, and infrastructure that helps the team move quickly
  • Work closely with research and infrastructure to ensure new models ship with optimized serving from day one

Systems intuition. Engineering depth. Appetite to go deep.

Nuance isn't looking for someone who's only fine-tuned models. They want someone who understands — or wants to deeply understand — the full stack from model weights to serving infrastructure.

  • 0–2 years of experience as an ML engineer (new grad level)
  • Hands-on experience profiling and optimizing LLM or diffusion models (NSight, torch profiler)
  • Experience productionizing ML models for inference and deployment at scale
  • Familiarity with vLLM, SGLang, or similar inference serving frameworks
  • Prior experience at a VC-backed startup or an AI team at top-tier big tech (Meta FAIR, Google DeepMind, Databricks) or strong LLM inference / ML systems research
  • BS, MS, or PhD in CS, ML, or a related field — completed or nearly done
  • Willing and able to work on-site in Seattle, 5 days a week
Python PyTorch CUDA vLLM SGLang TensorRT Triton Inference Server Rust Go Kubernetes Terraform WebRTC Ray

Small team. High bar. Full ownership.

Doing right by people

Integrity and respect aren't negotiable.

Transparency

Open communication, shared context — everyone can make informed decisions.

Relentless speed

Bias toward action. Iterate fast. Learn quickly.

In-person collaboration

The best ideas happen face-to-face. Five days a week in Seattle.

Fast, technical, and respectful of your time.

1
Intro Call
High-level conversation — your background, what you've built, and whether this is worth exploring further.
2
Technical Phone Screen 1
Deep-dive on ML inference, model serving, and systems design.
3
Technical Phone Screen 2
Harder questions on optimization approaches, specific stack experience, and how you'd tackle active Nuance challenges.
4
On-site Interview · Half Day in Seattle
Meet the team in person. Work through a technical problem together.
5
Offer
Move fast from on-site to offer for the right candidate.

Interested? Let's talk.

15 minutes is all it takes to find out if this is worth exploring further.

Book a Call with Shwetha