Job Openings Member of Technical Staff - Model Optimization and Inference

About the job Member of Technical Staff - Model Optimization and Inference

Member of Technical Staff - Model Optimization and Inference

Company: Nuance Labs
Location: Seattle, WA (in office 5 days per week; relocation assistance available)
Compensation: $250,000 - $350,000 + equity
Employment Type: Full-time
Visa Sponsorship: Visa sponsorship available

About Nuance Labs

Nuance Labs builds photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react and respond like a real person. The research team includes PhDs from MIT, UW, Oxford, CMU and Johns Hopkins.

Founded in 2024, Nuance Labs has raised $60M and has about 25 people.

The Role

Nuance Labs is hiring an experienced ML Infrastructure/Systems Engineer (2+ years) to own end-to-end inference optimization across LLMs, audio models and diffusion components, with a focus on latency, throughput and cost.

What You Will Do

  • Own end-to-end inference optimization across the model stack.
  • Implement and tune KV cache strategies for long-context conversations.
  • Evaluate, deploy and extend serving frameworks such as vLLM, SGLang and TensorRT-LLM.
  • Profile and benchmark latency and throughput and remove bottlenecks.
  • Accelerate diffusion inference and apply quantization techniques (INT8, INT4, GPTQ, AWQ).
  • Build internal tooling that makes optimization work faster and more rigorous.

What You Bring

  • 2+ years building and maintaining production ML systems
  • Designing scalable infrastructure from scratch
  • Track record optimizing latency, throughput and cost
  • Debugging distributed systems

Nice to Have

  • Video or audio model experience
  • CUDA kernels and low-level optimization
  • Real-time video streaming (WebRTC)

Benefits

HSA with about $2,000 annual company contribution, 15 days PTO plus public holidays and a company-wide office closure week.

Tech Stack

Kubernetes, Terraform, Python, Rust, Go, Dagster, Ray, Airflow, WebRTC, vLLM, Triton Inference Server, TensorRT