Valeria Wu (Google DeepMind) and Soham Ray (Sierra AI) explore the latest advances in native audio models and the challenges of building better real-time voice experiences. They discuss conversational latency, multilingual code-switching, agentic voice tasks, and how new benchmarks can better measure the quality of real-world dialogue.
Watch along and learn:
*How native speech-to-speech models preserve the tone and prosody traditional pipelines lose.
*The engineering tradeoffs behind sub-second response times and seamless, human-like interruption handling.
*How models smoothly navigate mixed-language phrasing and dialect shifts without dropping context.
*How to orchestrate API calls and stateful tasks while maintaining an uninterrupted vocal flow.
*Moving beyond Word Error Rate toward dynamic metrics that measure conversational flow and turn-taking.
Subscribe to Google for Developers →
Products Mentioned: Google DeepMind
Speakers: Valeria Wu, Soham Ray
Chapters:
00:00 – Introduction & The State of Real-Time Voice AI
0:50 – Understanding Sierra & TAU
01:48 – Latency
03:52 – Fluid Multilinguality
04:22 – What’s Next?
|
Download your free Python Cheat Sheet he...
Working With AI Agents: Short Live Cours...
Learn how to build useful AI agents that...
Download your free Python Cheat Sheet he...
How should you manage AI contributions t...
Learn how per-layer embeddings, or PLE, ...
From warehouse floors to boardrooms, AI ...
Valeria Wu (Google DeepMind) and Soham R...
Need to share QuickSight datasets and da...
Download your free Python Cheat Sheet he...
Building Flutter apps for desktop? Your ...