Confirm action

Are you sure you want to delete?

Link copied!
AI
Aug 27, 2026 · 2 min read

Gemini 3.5 Transcribe quietly becomes Google’s most accurate speech model

Affmarketingworld
Patric Mirgeschiss
Editor, Affmarketingworld
Gemini 3.5 Transcribe quietly becomes Google’s most accurate speech model

Google put Gemini 3.5 Transcribe into public preview with barely any fanfare, despite word-error rates that put it among the most accurate speech-to-text models the company has released.

What the numbers actually show

Google didn’t put together a launch event for this one. Gemini 3.5 Transcribe went into public preview on August 26 with a blog post and not much else, despite the model posting some of the lowest error rates the company has reported for speech-to-text. On Artificial Analysis benchmarks it lands at 4.0% word error rate in streaming mode and 2.6% for pre-recorded audio, and on the multilingual FLEURS benchmark those numbers come in at 5.50% and 5.04%. Google says latency also improved by 70% compared to Chirp 3, its previous transcription model. The company didn’t publish head-to-head numbers against OpenAI or Deepgram, so whatever ranking gets thrown around online right now is coming from outside testing, not from Google itself.

The model auto-detects whichever of 85-plus languages is being spoken, and it can switch languages mid-recording without needing to be told. Speaker separation works reliably for up to three people in a conversation; Google labels anything beyond that as experimental, so it’s not something to lean on yet for a crowded call. Every transcript comes with word-level timestamps, filler words get stripped out automatically, and the model handles mid-sentence corrections too — say “meet at 1, actually make it 2” and it just outputs the corrected version. There’s also support for custom vocabulary, which matters for anyone dealing with product names or technical jargon a general model would otherwise mangle. Access runs through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, with a slower rollout underway on Gboard for Android and the Gemini app on macOS.

A weekend project built on top of it

A Google developer used the same model to build Jot within days of the announcement, a free, open-source dictation tool for macOS 14 and up, released under Apache 2.0. Hold fn, talk, let go, and the cleaned-up text lands wherever the cursor already is — punctuation added, filler words gone, corrections applied automatically. Audio gets saved to disk right away so a crash or dead battery doesn’t wipe a dictation mid-sentence, and voice data goes straight from the Mac to Google’s API using the user’s own key, with no separate account or middleman server involved.

“Nobody staged a keynote for this one, and yet it’s putting up better numbers than models that got the full launch treatment.”

Patric Mirgeschiss
Reviewed by
Patric Mirgeschiss
Editor · AffMarketing World
Published Aug 27, 2026
X Profile →
Related tags