Gemini 3.8 Live tops speech benchmarks with 97-language support
Google launched Gemini 3.8 Live and its Extended Thinking variant on September 15, voice models built for agentic use with support for 97 languages and a top spot on speech-quality benchmarks.
97 languages, and thinking out loud
Google launched Gemini 3.8 Live and its Extended Thinking variant on September 15, both built specifically for voice-driven agentic use cases rather than text chat with a voice layer bolted on top. Language handling is the most immediately usable piece: the model automatically detects and switches between 97 supported languages mid-conversation, so a speaker can drop into a second language partway through and the model just follows without a restart or a manual toggle.
Extended Thinking changes how the model behaves under a harder question. Instead of going silent while it works, it reasons and speaks at the same time, filling the gap with something like “Let me check that” before narrating its own progress through a multi-step task as it runs in the background. That narration isn’t decoration. It’s covering the same latency a person would otherwise sit through as dead air, which is usually where a voice assistant starts feeling broken.
Tools mid-sentence, and a #1 benchmark spot
Both versions run tools and API calls without stopping the conversation to do it. The model can acknowledge a request and keep talking while the actual task finishes behind the scenes, rather than freezing mid-sentence the way most voice assistants still do when they need to fetch something external. Vision processing got fast enough that Google demoed it running a live chess game purely off visual input and turning a hand-drawn interface sketch into working code, both scenarios where any real lag breaks the illusion entirely.
The numbers back up the launch. Extended Thinking took the top spot on Artificial Analysis’ Speech to Speech Quality Index at 82.6, and scored 97.7% on Big Bench Audio. Every piece of generated audio carries an imperceptible SynthID watermark woven into the output itself, keeping AI-generated speech detectable even when a listener can’t hear anything unusual. Both models are live now through the Gemini API and Google AI Studio, with a rollout underway to Workspace, Search, and the Gemini app. This lands just days after Gemini 3.8 Flash shipped for coding and agentic workloads, which makes September a genuinely fast release cadence for a single model family rather than a one-off announcement.
“A voice model that narrates its own thinking instead of going silent sounds like a small trick, but dead air is exactly the moment every voice assistant has always sounded broken.”
Share
SUBSCRIBE TO OUR PRIVATE CASES AND USEFUL TIPS
Subscribe to our newsletter, get only exclusive content and weekly digests, no any spam!
By providing my email, I accept the Privacy Policy.