2026-09-22

Claude Opus 5.5: 'At the Level of Fable 5.1' Undersells It — the Table Shows a Clean Sweep, 9 for 9

AIModels🌍 North America

Anthropic launched Claude Opus 5.5: first model in the new Claude 5.5 family, priced at 4/4/20 per million tokens against Opus 5's 5/5/25, with cache reads cut to 0.20from0.20 from 0.50. The launch line: "performs at the level of Claude Fable 5.1 for most tasks."

Read the actual table, and that's an understatement. Opus 5.5 beats Fable 5.1 on all nine benchmarks shown — Terminal-Bench 4.0 (66.4% vs 55.8%), FrontierCode v1.1 (54.4% vs 50.3%), CursorBench 4.0 (57.8% vs 51.8%), GDPval-AA v2.1 (1846 vs 1735), AutomationBench (40.0% vs 31.4%), Humanity's Last Exam with tools (67.7% vs 65.6%), Terminal-Bench-Science (58.7% vs 52.6%), OSWorld 2.0 (81.8% vs 80.7%), and Chartography (89.0% vs 88.4%). Nine for nine, at 40% of Fable 5.1's price (4/4/20 vs 10/10/50). "At the level of" is the rare vendor claim that's conservative relative to its own data rather than inflated.

The two losses in the table come with a footnote explaining them

Against GPT-6 Astra specifically, Opus 5.5 loses exactly twice: AutomationBench (40.0% vs 41.4%) and Terminal-Bench-Science (58.7% vs 64.6%) — both shaded gray in the table, the only two rows marked that way. Anthropic's own footnotes address both directly. On safeguards generally: "When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." On AutomationBench specifically: "These runs were performed without fallback models, so safeguard interventions were considered failures — this resulted in a lower score than Claude Opus 5.5 would achieve in practice."

That's an unusual level of self-disclosed measurement noise — the only benchmarks Opus 5.5 loses are the ones its own footnotes flag as artificially depressed by how safety interventions get scored, not as genuine capability gaps. Worth taking at face value cautiously: "would achieve in practice" is Anthropic's own claim about itself, with no counterfactual run shown to back the number.

The reproduction-check footnotes are a smaller, cleaner positive: Anthropic re-ran the public Terminal-Bench 4.0 and Terminal-Bench-Science leaderboards for Claude Opus 5 as a sanity check and landed within noise both times (52.3% against the public 51.8%; 29.0% against 30.0%) — a good-faith methodology cross-check most launch tables skip.

At its cheapest setting, it already beats Sol's most expensive one

Terminal-Bench 4.0 score against cost per attempt, log scale, by effort level: Opus 5.5's curve (low, med, high, xhigh, max) sits above and to the left of Opus 5, Fable 5.1, GPT-6 Astra, and GPT-5.6 Sol at every point. Opus 5.5's cheapest setting, low, scores roughly 39% at about 1.30 — above GPT-5.6 Sol's most expensive point, roughly 37.5% at about 8. From Anthropic's Opus 5.5 launch page.

Anthropic's launch page adds the effort-level breakdown behind the single Terminal-Bench 4.0 number above (xhigh, 66.4%, matching this chart within rounding). Opus 5.5's cheapest setting — low, about 1.30perattemptscoresroughly391.30 per attempt — scores roughly 39%, already ahead of GPT-5.6 Sol's *most expensive* point on the chart, roughly 37.5% at about 8. The whole Opus 5.5 curve sits above and to the left of every rival at every cost level shown; it isn't Pareto-optimal at one point, it's Pareto-optimal across the entire range.

"40% less to run" is a blend, not any single price

No individual line item in the pricing table is a flat 40% cut: input tokens drop 20% (55→4), output tokens 20% (2525→20), cache writes 20% (6.256.25→5) — only cache reads drop that steeply, 60% (0.500.50→0.20). The 40% figure is a blended estimate across a representative workload mix, the same shape this blog flagged in Fable 5.1's own price cut: the actual savings on any given workload depends heavily on how cache-read-heavy it is, and a workload with little caching will see closer to 20% than 40%.