Google DeepMind shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking today, pitched as its best conversational AI yet: models that "talk, think, and handle tasks in the background without breaking your flow." The launch leans on outside benchmarks more than most model announcements this blog has covered — four separate charts from Artificial Analysis and one from Sierra, the enterprise-agent company. That's worth taking seriously, since this blog has noted before that Artificial Analysis is a genuinely independent benchmarking firm, not a house brand. But which chart a lab chooses to publish, and how it frames the number next to it, is still the lab's own editorial choice — and two details in Google's own selected charts undercut the framing around them.
The new flagship loses to last generation's high-effort mode
Start with the "Agentic Performance (τ-Voice)" chart, Google's own second bar chart in the launch materials. Gemini 3.8 Live Extended Thinking leads at 68.6%, ahead of OpenAI's GPT-Live-1 paired with GPT-6 Astra at 67.9%. Fine so far. But scan further down the same chart: Gemini 3.8 Live — the new, plain, non-Extended-Thinking model being launched today — scores 30.1%, below Gemini 3.1 Flash Live (High) at 37.7%. The previous generation's model, running at a higher effort setting, beats the new generation's base model on Sierra's own agentic-voice benchmark, by more than seven points. Nothing in Google's surrounding text acknowledges this. The blog post's prose about 3.8 Live focuses entirely on its "high preference among users" and cost efficiency — both real claims, addressed below — while the one chart that would let a reader compare the new base model against the old high-effort one shows the new model behind. A launch chart doesn't have to flatter every row to be honest, and Google didn't hide the number; it just didn't mention what it shows.
Two labs, the same benchmark family, two different winners
The "τ³-Banking Leaderboard" chart draws on Sierra's τ³-bench — a real, independently built benchmark suite covering retail, airline, telecom, banking, and voice tasks. Google's chart shows Gemini 3.8 Live Extended Thinking at 35.1%, ahead of GPT-Live-1 paired with GPT-6 Astra at 32.0%. That's a real result, on a real third-party benchmark. It also directly complicates a claim OpenAI made five days ago: when OpenAI launched GPT-Live-1 on September 10, it said the same GPT-Live-1-plus-Astra pairing "ranks first on Tau3" — the same Sierra benchmark family Google is now citing. Both claims can be true simultaneously if they're describing different subsets or configurations of a multi-domain suite (Tau3 spans more than banking), and this post can't fully reconcile the two from the public materials alone. But the pattern itself is familiar: this blog has now seen it with DeepSeek and Sakana's dueling Chartography numbers and again here — two labs citing the same named, real, third-party benchmark family, each publishing the specific cut where they come out ahead, neither publishing the other's.
"Second place" — behind whom?
Google's post claims Gemini 3.8 Live "has shown a high preference among users, securing a second place in the Speech Agent Arena." The Speech Agent Arena is Artificial Analysis's blind, live-conversation preference leaderboard — a real, independently run arena, not a Google construction. But Google's post doesn't say second place on which of the arena's two separate rankings: the Elo preference leaderboard, or the task-success-rate leaderboard, which aren't the same list. On the Elo preference leaderboard specifically, the model that has held the top spot in Artificial Analysis's own published rankings is Gemini 3.1 Flash Live Preview — Minimal, Google's own older, cheaper, previous-generation model, at 1,046 Elo. If "second place" refers to that leaderboard, the new 3.8 Live's own launch-day claim would mean it trails a cheaper model from Google's own prior generation, not a competitor's. That's not a knock on 3.8 Live's actual quality — arena rankings shift as new models accumulate votes, and 3.8 Live may well climb past its predecessor once enough conversations are logged — but it's a detail worth knowing rather than assuming "second place" means second to a rival lab.
What the cost chart actually shows
The one chart in the launch materials that reads cleanly in Google's favor is cost. Gemini 3.8 Live prices at $0.84 per hour of input audio on the Big Bench Audio subset — the cheapest model on the chart, undercutting even last generation's Flash Live variants ($1.50–$1.75) and far below GPT-Live-1 Astra ($5.83) or Grok Voice Think Fast 2.0 ($4.80). Extended Thinking costs more than four times as much as the base model, at $3.50 — still cheaper than the two named competitors, but a real jump from the base tier for whatever the extra reasoning buys. Read next to the agentic-performance chart above, the honest read is that Google shipped a genuinely cheap conversational model and a genuinely more capable but much pricier reasoning model, and the "3.8 Live" name covers both — worth knowing which one a given benchmark chart is actually describing before assuming "3.8 Live" means the same thing in every bar.
SynthID, extended to voice
Google says all audio from these models is watermarked with SynthID. This blog wrote a full technical explainer on SynthID-Text last month — the mechanism Google DeepMind published in Nature, which Anthropic later confirmed its own Claude watermark is a version of. The core idea (replacing a model's random token choices with a secretly-keyed pseudo-random source, so the same output distribution becomes recognizable later without changing what the model says) was designed for text, but the same logic extends naturally to audio generation, where a model is still making a long sequence of probabilistic choices. Google's launch post doesn't say whether Gemini's audio watermark uses the identical SynthID mechanism or a modality-specific variant, which would be worth confirming before assuming the text explainer's caveats — low-entropy content carries little signal, paraphrasing erodes it — transfer directly to spoken audio.
A benchmark from a company on the partner-quote list
One more chart worth a smaller note: EVA-Bench, which Google says its models use to "push the Pareto Frontier for complex workflows," is ServiceNow's own benchmark, run, per Google's own footnote, "on the Live API on Gemini Enterprise Agent Platform." ServiceNow also appears among the partner quotes in the same launch post, alongside Salesforce, Genspark, and Lumeris. That doesn't make EVA-Bench's result false, but it's a data point from a commercial partner testing on Google's own platform, not an outside lab's independent evaluation — a different category of evidence than the Artificial Analysis and Sierra charts, worth not conflating with them even though all five charts sit in the same post with the same visual weight.
Where this leaves the launch
None of this amounts to Gemini 3.8 Live being a bad model — the Speech to Speech Index chart, the clearest single measure of raw conversational quality, has Extended Thinking genuinely leading a real field, and the cost numbers are a real, checkable advantage. What it shows is the now-familiar pattern of a frontier lab's own launch materials containing the evidence for a more complicated story than the surrounding prose tells: a new base model that regresses on the one agentic benchmark shown, a benchmark-family claim that overlaps uncomfortably with a rival's five-day-old claim on the same suite, and a "second place" line whose meaning depends on a ranking the post never names. Reading the charts themselves, not just the paragraphs next to them, is doing real work here.