Top 3 new model launches at Gemini Audio at Night
Google DeepMind releases new audio models including TTS, live translation, and Gemini Live upgrades
“Our new text-to-speech models are capable of doing really compelling voice personalization, which means that you can create a voice and guide it to have specific emotions, specific resonance, and even to incorporate things like pauses.”
Google DeepMind announced three new audio capabilities at Google I/O: expressive text-to-speech with emotion and resonance control, live translation across 100 languages from video or audio input, and expanded Gemini Live models with function/tool call support for developers. These updates extend Gemini's multimodal audio stack and lower the barrier for developers building voice-native AI applications.