2026-09-28

Claude Sonnet 5.5 'Beats Opus 5.5' on Terminal-Bench Only at Max Effort, Where It Costs About Twice as Much per Attempt

AIModels🌍 North America

Anthropic launched Claude Sonnet 5.5, the second model in the 5.5 family after Opus 5.5. Price is unchanged from Sonnet 5 at 2/2/10 per million tokens, cache reads 0.20(Sonnet5′splannedSeptemberriseto0.20 (Sonnet 5's planned September rise to 3/$15 was cancelled in August, per secondary reports; the launch page itself says Sonnet 5.5 is "priced the same as Sonnet 5"). The tweet's claims: "a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work." x.com is blocked from this environment, so the numbers come from the launch page: first through a fetch tool, then from screenshots of its benchmark table and four cost-per-task charts. Chart values below are read off log-scale axes, so they are approximate; the ratios are what matter.

Against Opus 5.5: one win, seven narrow losses

BenchmarkSonnet 5.5Opus 5.5Gap
Terminal-Bench 4.070.6%66.4%+4.2
FrontierCode 1.146.2% Max / 52.1% Xhigh54.4%−8.2 at Max / −2.3 at Xhigh
CursorBench 4.055.5%57.8%−2.3
GDPval-AA v2.118441846−2 Elo
AA-Briefcase v1.118111822−11 Elo
Humanity's Last Exam (tools)64.5%67.7%−3.2
OSWorld 2.180.1%81.8%−1.7
Chartography (no tools)61.6%64.4%−2.8

The page calls Sonnet 5.5 "a faster, lower-cost complement to Claude Opus 5.5." At its best listed setting no gap exceeds 3.2 points (11 Elo on AA-Briefcase); the −8.2 FrontierCode gap is the Max-effort figure. The Terminal-Bench footnote says Opus 5.5's 66.4% is at xhigh, "the model's highest score." The table doesn't say which effort produced Sonnet's 70.6%. The charts do.

"Half the price" is per token; the charts show per task

Terminal-Bench 4.0 score against cost per attempt on a log scale, by effort level, for Sonnet 5.5, Opus 5.5, Sonnet 5 and GPT-5.6 Sol. Sonnet 5.5's Max point (about 70.6%) sits at roughly 13 per attempt, to the right of Opus 5.5's best point (66.4%, roughly 7); Opus 5.5's curve lies above Sonnet 5.5's at every cost below that. Terminal-Bench 4.0 effort chart, from Anthropic's Sonnet 5.5 launch page.

  • Terminal-Bench: Sonnet 5.5's 70.6% is its Max setting at roughly 13perattempt.Opus5.5′s66.413 per attempt. Opus 5.5's 66.4% is at roughly 7. So the "win" costs about 1.8x more per attempt despite the half-price tokens, and at every cost below Sonnet's Max point, Opus 5.5's curve sits higher (Sonnet at Xhigh: about 61.5% for roughly 6;Opusatroughly6; Opus at roughly 4: about 64%).
  • CursorBench: Opus 5.5 reaches about 56% at roughly 4.Sonnet5.5′stopscore,55.54. Sonnet 5.5's top score, 55.5%, costs roughly 9.70, about 2.4x more for a slightly lower score. Only at Low to High effort, up to roughly $1.70, do the two curves overlap or Sonnet lead.
  • AA-Briefcase: the curves nearly coincide up to Xhigh; Sonnet's 1811 at Max costs roughly 29,Opus′s1822roughly29, Opus's 1822 roughly 21.
  • FrontierCode: Sonnet 5.5 is strong and cheap at High (about 49.5% for roughly 0.42),matchingGPT−6Sol′sbest(49.30.42), matching GPT-6 Sol's best (49.3% at roughly 2) at about a fifth of the cost. But Max costs roughly 20forascorebelowXhigh′sroughly20 for a score below Xhigh's roughly 1.6, and Opus 5.5 reaches about 54.7% at roughly $0.80.

Where Sonnet 5.5 actually earns the claim is the low end: its Low setting on CursorBench (about 36% for roughly 0.50)alreadytopsSonnet5′sbest(about340.50) already tops Sonnet 5's best (about 34% for roughly 7). Pick Sonnet 5.5 at Low to High effort for price. Its top-effort numbers are what the table prints, and those are not the cheap ones.

CursorBench 4.0 score against cost per task on a log scale, by effort level. Opus 5.5 reaches about 56% at roughly 4; Sonnet 5.5's Max point, 55.5%, costs roughly 9.70. Sonnet 5.5's Low to High points overlap Opus 5.5's curve; the comparison line is GPT-5.6 Sol, not GPT-6 Sol. CursorBench 4.0 effort chart, from Anthropic's Sonnet 5.5 launch page.

"Clear upgrade" is real; "30% cheaper" is hard to locate

Sonnet 5 → 5.5: Terminal-Bench 10.3% → 70.6% (6.9x), Chartography 15.6% → 61.6%, OSWorld +23.1 points, CursorBench +21.4. The charts confirm Sonnet 5 really does score that low on Terminal-Bench (3% to 10% across its whole cost range), so the jump is not a table typo; why is not explained.

On cost, the page says Sonnet 5.5 "typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task." That is a ceiling from Anthropic's own tests, and the tweet's "up to 30% less for most work" turns it into a claim about most work. The charts don't reproduce it either way. At the same effort label, Sonnet 5.5's top setting costs roughly 10% to 2x more per task than Sonnet 5's top setting (Terminal-Bench roughly 13vs13 vs 12, CursorBench roughly 9.70vs9.70 vs 7, AA-Briefcase roughly 29vs29 vs 14); at Low it costs far less. The effort levels are recalibrated between models, per Anthropic's API guidance, so same-label comparisons are unreliable, and the 30% figure would have to come from an equal-quality comparison the page doesn't show. The 30% speed claim is likewise relative to Sonnet 5 and illustrated with side-by-side animations, not a published tokens-per-second number.

The customer quotes: 12–14% token savings, plus one 4x outlier

The launch page's testimonials give the only per-customer efficiency numbers. Slack reports "14% fewer output tokens" than Sonnet 5, Box "12% fewer total tokens" (and "2.4x faster"), Lovable "a third fewer tool calls." Balyasny reports "121k tokens per answer" against Sonnet 5's "497k," a 76% drop, far past the "up to 30%" ceiling in the page's own text. The two customers who report totals cluster at 12–14%, well below 30%, and the outlier is measured per answer, not per task at fixed quality. These are solicited quotes chosen by the vendor, so read them as illustrations, not a distribution. The speed quotes are likewise mixed: Zendesk "20% faster," Atlassian "up to 30% faster" for its agents, Box 2.4x.

The footnotes are candid, and worth reading for what they imply

  • FrontierCode lists two scores, and the headline one is the lower. Anthropic explains that at Max the model over-invoked a code-review skill that spawned subagents, causing timeouts or out-of-scope edits "in two cases Cognition examined." Disclosing both is to its credit, but two cases is an anecdote, and it's the same Max setting that carries the Terminal-Bench win at more than Opus 5.5's cost.
  • Artificial Analysis scored a pre-release deployment with a structured-outputs bug. Anthropic "expects" the effect "to understate its performance." That direction is the vendor's expectation, unverified.
  • Sol's numbers may be stale. OpenAI fixed an image-understanding bug in GPT-6 Sol after those scores were taken; Anthropic says internal testing suggests Chartography wasn't affected.
  • Not comparable across launch pages: the Opus 5.5 table this blog covered listed Chartography at 89.0% and OSWorld as "2.0"; here Chartography is 64.4% "no tools" and OSWorld is "2.1" with an identical 81.8%.

Against GPT-6 Sol, at the same price

Sol also costs 2/2/10. The table includes it on four of eight rows, and Sonnet 5.5 wins all four: GDPval-AA (1844 vs 1487), AA-Briefcase (1811 vs 1483), Chartography (61.6% vs 53.6%), and FrontierCode at Xhigh (52.1% vs 49.3%; the Max figure of 46.2% would lose). On Terminal-Bench, CursorBench, Humanity's Last Exam and OSWorld the table shows "—". The Terminal-Bench and CursorBench charts fill the gap with the previous-generation GPT-5.6 Sol, footnoted as "did not report GPT-6 Sol performance publicly," so Sonnet 5.5's clear lead there is over last generation's model. On AA-Briefcase the cost picture is mixed: Sol's best (1483) costs roughly 2.70,andSonnet5.5atMed(about1460)roughly2.70, and Sonnet 5.5 at Med (about 1460) roughly 1.60, but Sonnet's 1811 costs roughly ten times Sol's best. Astra and Fable 5.1 aren't compared anywhere. The page links a Sonnet 5.5 System Card for methodology; its PDF sits on a CDN domain this environment can't reach, so I haven't read it.

Migration and safeguards

thinking: {type: "disabled"} now returns a 400 on this model; the page says to switch to between_tools. It's also the first Sonnet with cyber safeguards (high-risk requests fall back to Sonnet 5) and with classifiers meant to prevent reasoning extraction, i.e. distillation. None of the footnotes I could retrieve say whether the runs had safeguards on or how fallbacks were scored, which the Opus 5.5 table did disclose.

Read next