Anthropic launched Claude Opus 5.5: first model in the new Claude 5.5 family, priced at 20 per million tokens against Opus 5's 25, with cache reads cut to 0.50. The launch line: "performs at the level of Claude Fable 5.1 for most tasks."
Read the actual table, and that's an understatement. Opus 5.5 beats Fable 5.1 on all nine benchmarks shown — Terminal-Bench 4.0 (66.4% vs 55.8%), FrontierCode v1.1 (54.4% vs 50.3%), CursorBench 4.0 (57.8% vs 51.8%), GDPval-AA v2.1 (1846 vs 1735), AutomationBench (40.0% vs 31.4%), Humanity's Last Exam with tools (67.7% vs 65.6%), Terminal-Bench-Science (58.7% vs 52.6%), OSWorld 2.0 (81.8% vs 80.7%), and Chartography (89.0% vs 88.4%). Nine for nine, at 40% of Fable 5.1's price (20 vs 50). "At the level of" is the rare vendor claim that's conservative relative to its own data rather than inflated.
The two losses in the table come with a footnote explaining them
Against GPT-6 Astra specifically, Opus 5.5 loses exactly twice: AutomationBench (40.0% vs 41.4%) and Terminal-Bench-Science (58.7% vs 64.6%) — both shaded gray in the table, the only two rows marked that way. Anthropic's own footnotes address both directly. On safeguards generally: "When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." On AutomationBench specifically: "These runs were performed without fallback models, so safeguard interventions were considered failures — this resulted in a lower score than Claude Opus 5.5 would achieve in practice."
That's an unusual level of self-disclosed measurement noise — the only benchmarks Opus 5.5 loses are the ones its own footnotes flag as artificially depressed by how safety interventions get scored, not as genuine capability gaps. Worth taking at face value cautiously: "would achieve in practice" is Anthropic's own claim about itself, with no counterfactual run shown to back the number.
The reproduction-check footnotes are a smaller, cleaner positive: Anthropic re-ran the public Terminal-Bench 4.0 and Terminal-Bench-Science leaderboards for Claude Opus 5 as a sanity check and landed within noise both times (52.3% against the public 51.8%; 29.0% against 30.0%) — a good-faith methodology cross-check most launch tables skip.
At its cheapest setting, it already beats Sol's most expensive one
From Anthropic's Opus 5.5 launch page.
Anthropic's launch page adds the effort-level breakdown behind the single Terminal-Bench 4.0 number above (xhigh, 66.4%, matching this chart within rounding). Opus 5.5's cheapest setting — low, about 8. The whole Opus 5.5 curve sits above and to the left of every rival at every cost level shown; it isn't Pareto-optimal at one point, it's Pareto-optimal across the entire range.
"40% less to run" is a blend, not any single price
No individual line item in the pricing table is a flat 40% cut: input tokens drop 20% (4), output tokens 20% (20), cache writes 20% (5) — only cache reads drop that steeply, 60% (0.20). The 40% figure is a blended estimate across a representative workload mix, the same shape this blog flagged in Fable 5.1's own price cut: the actual savings on any given workload depends heavily on how cache-read-heavy it is, and a workload with little caching will see closer to 20% than 40%.