Google DeepMind announced Gemini 3.5 Transcribe, a speech-to-text model succeeding Chirp 3, shipped across two APIs: real-time streaming with sub-second latency via the Live API (gemini-3.5-transcribe-live), and pre-recorded audio processing with speaker attribution and word-level timestamps via the Interactions API (gemini-3.5-transcribe). The feature list is straightforward and genuinely useful for a transcription model: handling mid-sentence self-corrections ("let's meet Tuesday — no, Wednesday"), stripping filler words, auto-formatting output, recognizing custom vocabulary, and auto-detecting more than 85 languages. Speaker attribution is explicitly capped and disclosed as such: up to three speakers supported, with anything beyond that marked "experimental" rather than silently degraded.
The headline number, and where it comes from
Google states the model "achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases," attributed directly to Artificial Analysis — a real, independent benchmark aggregator, not a number Google generated itself. The same source is credited for a separate claim: time to final transcription improves by 70% over Chirp 3. Both are genuinely third-party-measured figures, worth taking at face value as real, checkable claims rather than vendor self-reporting.
The number that doesn't match, on the one benchmark actually named
Further down the same post, Google reports results on FLEURS — a standard, publicly known multilingual speech benchmark — "across a set of top languages and locales": 5.50% WER in streaming mode and 5.04% WER in non-streaming. Worth sitting with the gap directly: the non-streaming FLEURS figure, 5.04%, is nearly double the 2.6% headline non-streaming average reported just above it in the same announcement. That's not necessarily a contradiction — an "average" across whatever mix of languages and conditions Artificial Analysis tested is a different, and likely easier, distribution than a specifically multilingual benchmark spanning many languages and locales at once — but it's a real, specific gap between the number Google leads with and the number tied to the one named, externally recognizable benchmark in the post. Worth noting too: unlike the headline WER and the 70%-latency figures, the FLEURS numbers aren't explicitly attributed to Artificial Analysis or any other named third party in the text — they read as Google's own internal evaluation, a different evidentiary tier than the numbers measured externally just above them. And despite repeating "improving over Chirp 3" twice, the announcement never states what Chirp 3's own WER or FLEURS score actually was — the improvement claim has no baseline attached anywhere in the post to check it against.
Function calling, and where it actually works today
One capability worth flagging for what it currently is rather than what it sounds like: the model "can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls" — a real, interesting design where a transcription model routes work to other models based on voice commands. The post's own text specifies this is "currently available in the Gemini macOS app" only, not the general API. That's a meaningful scope limit on an otherwise broad-sounding capability claim, worth keeping attached to it.
The rollout is staged, not uniform
"Start using it today" undersells how partial availability actually is at launch. For developers: public preview in Google AI Studio and Google Antigravity. For enterprises: public preview via the Gemini Enterprise Agent Platform, with the Customer Experience product line explicitly "coming soon" rather than live. For consumers: the Gemini app on macOS (English only), the Rambler feature on Android (select countries and languages only), and Chrome support that's "coming soon" and not live at all yet. Three different products, three different availability states, one launch date — worth reading the specific surface before assuming universal access.
What's a customer testimonial, and what's an independent integration
Google names Vivo, Intellitek Health, and Lingopal as companies giving positive feedback — standard vendor-selected testimonials, worth treating as marketing quotes rather than independent evaluation. Separately, and more substantively: real third-party developer platforms — Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents — are named as already building real-time voice infrastructure on top of the Live API. That's a different, more checkable kind of claim than a testimonial: these are actual products a reader could go verify are integrating the API, not quotes selected for the announcement.
What to expect next
- Watch for a direct, same-source WER comparison between the average number and the FLEURS number. Right now the two headline accuracy claims in the same post use different baselines and, apparently, different measurement sources — worth Google clarifying rather than leaving the gap unexplained.
- Watch for Chirp 3's actual baseline numbers. Every improvement claim in the launch post is relative ("improves over Chirp 3," "improves by 70%") without stating what Chirp 3 itself scored — independently comparable numbers for both models would make the improvement claims checkable rather than just asserted.
- Watch the staged rollout resolve. Chrome support, wider consumer-language coverage, and the Enterprise Customer Experience product are all explicitly not live yet — worth checking back once "coming soon" becomes shipped.
References: Google — Gemini 3.5 Transcribe announcement · related coverage: Gemini 3.7 Flash Wins Nine of Twenty Benchmarks · Frontier Arcade: trends & predictions