Anthropic launched Claude Sonnet 5.5, the second model in the 5.5 family after Opus 5.5. Price is unchanged from Sonnet 5 at 10 per million tokens, cache reads 3/$15 was cancelled in August, per secondary reports; the launch page itself says Sonnet 5.5 is "priced the same as Sonnet 5"). The tweet's claims: "a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work." x.com is blocked from this environment, so the numbers come from the launch page: first through a fetch tool, then from screenshots of its benchmark table and four cost-per-task charts. Chart values below are read off log-scale axes, so they are approximate; the ratios are what matter.
Against Opus 5.5: one win, seven narrow losses
| Benchmark | Sonnet 5.5 | Opus 5.5 | Gap |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | +4.2 |
| FrontierCode 1.1 | 46.2% Max / 52.1% Xhigh | 54.4% | −8.2 at Max / −2.3 at Xhigh |
| CursorBench 4.0 | 55.5% | 57.8% | −2.3 |
| GDPval-AA v2.1 | 1844 | 1846 | −2 Elo |
| AA-Briefcase v1.1 | 1811 | 1822 | −11 Elo |
| Humanity's Last Exam (tools) | 64.5% | 67.7% | −3.2 |
| OSWorld 2.1 | 80.1% | 81.8% | −1.7 |
| Chartography (no tools) | 61.6% | 64.4% | −2.8 |
The page calls Sonnet 5.5 "a faster, lower-cost complement to Claude Opus 5.5." At its best listed setting no gap exceeds 3.2 points (11 Elo on AA-Briefcase); the −8.2 FrontierCode gap is the Max-effort figure. The Terminal-Bench footnote says Opus 5.5's 66.4% is at xhigh, "the model's highest score." The table doesn't say which effort produced Sonnet's 70.6%. The charts do.
"Half the price" is per token; the charts show per task
Terminal-Bench 4.0 effort chart, from Anthropic's Sonnet 5.5 launch page.
- Terminal-Bench: Sonnet 5.5's 70.6% is its Max setting at roughly 7. So the "win" costs about 1.8x more per attempt despite the half-price tokens, and at every cost below Sonnet's Max point, Opus 5.5's curve sits higher (Sonnet at Xhigh: about 61.5% for roughly 4: about 64%).
- CursorBench: Opus 5.5 reaches about 56% at roughly 9.70, about 2.4x more for a slightly lower score. Only at Low to High effort, up to roughly $1.70, do the two curves overlap or Sonnet lead.
- AA-Briefcase: the curves nearly coincide up to Xhigh; Sonnet's 1811 at Max costs roughly 21.
- FrontierCode: Sonnet 5.5 is strong and cheap at High (about 49.5% for roughly 2) at about a fifth of the cost. But Max costs roughly 1.6, and Opus 5.5 reaches about 54.7% at roughly $0.80.
Where Sonnet 5.5 actually earns the claim is the low end: its Low setting on CursorBench (about 36% for roughly 7). Pick Sonnet 5.5 at Low to High effort for price. Its top-effort numbers are what the table prints, and those are not the cheap ones.
CursorBench 4.0 effort chart, from Anthropic's Sonnet 5.5 launch page.
"Clear upgrade" is real; "30% cheaper" is hard to locate
Sonnet 5 → 5.5: Terminal-Bench 10.3% → 70.6% (6.9x), Chartography 15.6% → 61.6%, OSWorld +23.1 points, CursorBench +21.4. The charts confirm Sonnet 5 really does score that low on Terminal-Bench (3% to 10% across its whole cost range), so the jump is not a table typo; why is not explained.
On cost, the page says Sonnet 5.5 "typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task." That is a ceiling from Anthropic's own tests, and the tweet's "up to 30% less for most work" turns it into a claim about most work. The charts don't reproduce it either way. At the same effort label, Sonnet 5.5's top setting costs roughly 10% to 2x more per task than Sonnet 5's top setting (Terminal-Bench roughly 12, CursorBench roughly 7, AA-Briefcase roughly 14); at Low it costs far less. The effort levels are recalibrated between models, per Anthropic's API guidance, so same-label comparisons are unreliable, and the 30% figure would have to come from an equal-quality comparison the page doesn't show. The 30% speed claim is likewise relative to Sonnet 5 and illustrated with side-by-side animations, not a published tokens-per-second number.
The customer quotes: 12–14% token savings, plus one 4x outlier
The launch page's testimonials give the only per-customer efficiency numbers. Slack reports "14% fewer output tokens" than Sonnet 5, Box "12% fewer total tokens" (and "2.4x faster"), Lovable "a third fewer tool calls." Balyasny reports "121k tokens per answer" against Sonnet 5's "497k," a 76% drop, far past the "up to 30%" ceiling in the page's own text. The two customers who report totals cluster at 12–14%, well below 30%, and the outlier is measured per answer, not per task at fixed quality. These are solicited quotes chosen by the vendor, so read them as illustrations, not a distribution. The speed quotes are likewise mixed: Zendesk "20% faster," Atlassian "up to 30% faster" for its agents, Box 2.4x.
The footnotes are candid, and worth reading for what they imply
- FrontierCode lists two scores, and the headline one is the lower. Anthropic explains that at Max the model over-invoked a code-review skill that spawned subagents, causing timeouts or out-of-scope edits "in two cases Cognition examined." Disclosing both is to its credit, but two cases is an anecdote, and it's the same Max setting that carries the Terminal-Bench win at more than Opus 5.5's cost.
- Artificial Analysis scored a pre-release deployment with a structured-outputs bug. Anthropic "expects" the effect "to understate its performance." That direction is the vendor's expectation, unverified.
- Sol's numbers may be stale. OpenAI fixed an image-understanding bug in GPT-6 Sol after those scores were taken; Anthropic says internal testing suggests Chartography wasn't affected.
- Not comparable across launch pages: the Opus 5.5 table this blog covered listed Chartography at 89.0% and OSWorld as "2.0"; here Chartography is 64.4% "no tools" and OSWorld is "2.1" with an identical 81.8%.
Against GPT-6 Sol, at the same price
Sol also costs 10. The table includes it on four of eight rows, and Sonnet 5.5 wins all four: GDPval-AA (1844 vs 1487), AA-Briefcase (1811 vs 1483), Chartography (61.6% vs 53.6%), and FrontierCode at Xhigh (52.1% vs 49.3%; the Max figure of 46.2% would lose). On Terminal-Bench, CursorBench, Humanity's Last Exam and OSWorld the table shows "—". The Terminal-Bench and CursorBench charts fill the gap with the previous-generation GPT-5.6 Sol, footnoted as "did not report GPT-6 Sol performance publicly," so Sonnet 5.5's clear lead there is over last generation's model. On AA-Briefcase the cost picture is mixed: Sol's best (1483) costs roughly 1.60, but Sonnet's 1811 costs roughly ten times Sol's best. Astra and Fable 5.1 aren't compared anywhere. The page links a Sonnet 5.5 System Card for methodology; its PDF sits on a CDN domain this environment can't reach, so I haven't read it.
Migration and safeguards
thinking: {type: "disabled"} now returns a 400 on this model; the page says to switch to between_tools. It's also the first Sonnet with cyber safeguards (high-risk requests fall back to Sonnet 5) and with classifiers meant to prevent reasoning extraction, i.e. distillation. None of the footnotes I could retrieve say whether the runs had safeguards on or how fallbacks were scored, which the Opus 5.5 table did disclose.