Anthropic released Claude Opus 5 on July 24, 2026, and the official pitch is deliberately modest: nearly all the intelligence of Fable 5 at half the cost β 25 per million tokens against Fable's 50 β sitting, in Anthropic's own tier system, a class below its Mythos-line flagship.
Then the independent numbers came in, and the modest pitch stopped holding. Opus 5 debuted at #1 on the Artificial Analysis Intelligence Index β 61, edging out Fable 5's 60 and pushing GPT-5.6 Sol to third. On Anthropic's own benchmarks it beats Fable 5 on coding and knowledge work: 43.3% on Frontier-Bench (agentic terminal coding), more than double Opus 4.8's 18.7% and well clear of Fable 5's 33.7%; within half a point of Fable 5's best on CursorBench at roughly half the cost per task; ahead of everything at any given price point on OSWorld computer use.
ARC Prize's own verified leaderboards make the case even more starkly. On ARC-AGI-1, Opus 5 scores 97.5% at 2.06/task β both, per ARC Prize, "competitive with previous SOTA performance for slightly higher cost." Solid, not shocking. But ARC-AGI-3 β the newest benchmark, built around interactive, agentic puzzle environments rather than static ones β is where the gap turns genuinely strange: Opus 5 sets a new SOTA at 30.2%, against a previous best of just 7.8% set by GPT-5.6 Sol (Max). Opus 5 isn't leading by a margin, it's nearly quadrupling the previous ceiling. The understudy is, by several measures, now the lead β while officially remaining a tier below.
Turning a puzzle into an equation
The most striking detail in ARC Prize's writeup isn't the score β it's how Opus 5 got there. ARC-AGI-3 environments are visual, interactive puzzles, not math problems. But partway through its transcript, on action 23, the model described a scene using explicit algebraic notation: 4_center = 2Γaxis β 5_center β a reflection equation, derived on its own, to represent what it was seeing. ARC Prize calls it the first time they've observed a model produce an explicit reflection equation in this kind of analysis.
It didn't stop there. By action 248, Opus 5 had generalized the pattern to two dimensions: "ghost = reflection about axis strips: br' = 2Β·br_axis β br, bc' = 2Β·bc_axis β bc." That's a model spontaneously discovering and formalizing a general geometric transformation rule mid-task, rather than pattern-matching against training data β exactly the kind of behavior ARC-AGI-3 was built to surface and that no prior frontier model had shown on it.
The features that matter
The 1M-token context window confirms the pre-launch leak, and the headline product feature is effort controls: low/medium/high plus a new "xhigh" reasoning mode, adjustable per turn β the cost/capability dial moving from an infrastructure detail to a first-class user control. It shipped everywhere at once: API, the Claude apps, Claude Code, Bedrock, Vertex, and GitHub Copilot same-day, as the new default on Claude Max.
The system card has an interesting inversion: Opus 5 tested less capable at cyber-exploitation than Fable 5, so it ships with lighter safeguards β and unlike Fable 5, it doesn't carry the 30-day data-retention regime Fable acquired during its export-control saga. Anthropic calls it "the most aligned Opus model" and its most capable generally available model for scientific research.
The prediction ledger, again
The trends article flagged Anthropic as overdue on its release cadence β projected July 12 β and argued, after the FLUX 3 hit, that silence past a lab's projected date tends to precede a bigger release. Opus 5 landed July 24, twelve days past projection: a genuinely big release, arriving late, exactly on pattern. That's two for two for the cadence math in one week β first Black Forest Labs, now Anthropic. (Google's mid-August projection is next on the board.)
What the launch conspicuously doesn't say
Two days after the White House accused Moonshot of distilling Anthropic's Fable, and three days before Kimi K3's open weights are due, Anthropic's launch materials say nothing about distillation, weight security, or the restriction fight in which it is the alleged victim and presumptive beneficiary. The closest thing to a comment came from a product lead who, asked about K3, offered only that it "remains to be seen" how open-weight models perform on complicated real-world projects. Whether that silence is legal caution, strategic patience, or discomfort with a ban being argued in its name, it's the loudest quiet part of an otherwise very loud launch β and the pricing tells its own story: cutting the effective cost of near-Fable intelligence in half, three days before the biggest open-weight release of the year, is a competitive answer that doesn't require saying anything at all.
Links
- Announcement: Claude Opus 5 β Anthropic
- Verified scores: ARC Prize β Claude Opus 5 results