2026-08-12

Grok 4.6: 'Matches GPT-5.6 Sol' Is True, and Also the Least Interesting Line in the Table

AIBenchmarks🌍 North America

xAI shipped Grok 4.6 today, a day after Grok Bot's public beta launch — the same shipping cadence we flagged as the point of the last four months of SpaceXAI/Cursor dealmaking. The positioning is explicit: "a particular focus on long-running agents and more ambitious interactive and visual work," staying with complex multi-step tasks — research, codebase work, or turning an idea into a finished artifact. xAI's own headline claim is that Grok 4.6 "matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index." That's true. It's also the least interesting sentence in a nine-benchmark table that tells a more specific story than the composite score does.

What the table actually shows

Here is xAI's own comparison table, reproduced: Grok 4.6 High vs. Grok 4.5 High vs. GPT-5.6 Sol Max vs. Claude Fable 5 Max. One caveat up front: xAI describes the third-party figures as "the best of self-reported or publicly available results," which means xAI chose which of its rivals' own numbers to cite — not that xAI ran the comparisons itself.

BenchmarkGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%—58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Count the wins and the composite-score framing gets less flattering: Fable 5 Max leads on five of the nine metrics (AA Intelligence Index, CursorBench, FrontierCode, APEX-Agents, APEX-SWE), GPT-5.6 Sol Max leads on two (DeepSWE, Terminal-Bench), and Grok 4.6 leads on three (GDPVal-AA, AA-Briefcase, Harvey LAB). "Matches GPT-5.6 Sol" is accurate on the one composite metric xAI chose to headline. But it's not the metric where Grok 4.6 actually looks strongest, and it's not close to describing Fable 5's position — Fable 5 leads more often than either of the models xAI compared itself to directly.

The real pattern: knowledge work and one specific vertical, not classic coding agents

The shape of where Grok 4.6 actually wins is more informative than the headline. Its three clean leads — GDPVal-AA (a real-world economic-value task index), AA-Briefcase (knowledge-work tasks), and Harvey LAB (Vals) — cluster around broad office and professional work. The standout is Harvey LAB specifically: a legal-domain benchmark where Grok 4.6 posts 15.8% against GPT-5.6 Sol's startlingly low 2.5%. That's a 6x gap on one specific vertical, which reads either as a genuine specialization or as a benchmark GPT-5.6 Sol simply wasn't tuned for — worth independent scrutiny either way.

Where it doesn't lead is the classic coding-agent cluster: DeepSWE, Terminal-Bench, CursorBench, FrontierCode, and APEX-SWE all go to either Fable 5 or GPT-5.6 Sol. That sits in some tension with the announcement's own framing, which explicitly names "working across a codebase" as a target use case. The numbers say Grok 4.6 is a competent-but-not-leading coding agent and a genuinely strong knowledge-work and legal-domain model. The interesting Grok 4.6 story is narrower and more specific than "frontier intelligence" implies, and also more useful if it holds up: a model that's good at office and professional tasks isn't a lesser story than a coding benchmark leader, it's a different one.

Training: the model teaching its own successor's data

xAI describes a longer supplemental training run than 4.5's, with curated model-generated data for reasoning and technical concepts, plus an improved optimizer and recipe as the SFT/RL foundation. The more interesting detail: Grok 4.5 was used to regenerate the SFT trajectories for 4.6 — across reasoning efforts, agent harnesses, and domains — with model-based filtering to drop bad traces. That's the same self-bootstrapping pattern we've now seen from Meta (Spark 1.1 grading Spark 1.2's candidate solutions) and implicitly from Liquid's MOPD: use the previous generation as a teacher for the next generation's training data, rather than only generating from scratch. It's becoming a standard move across labs, not a one-off technique.

RL training covered "a wide range of agentic RL tasks" — knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and CAD. xAI also reports Grok 4.6 showing more self-testing and verification behavior on longer trajectories — "checking its own work before moving on" — which is a qualitative, unbenchmarked claim, not a scored metric, and worth treating that way until someone measures it directly.

The safety section says a lot by saying little

Compare the specificity gap. On one side: nine quantified capability benchmarks with exact scores. On the other: a safety paragraph that says safeguards were "improved and calibrated in line with the model's capabilities," backed by "our widest-ever suite of pre-deployment testing" and "extensive post-deployment and third-party testing" — no numbers, no named third parties, no published red-team results alongside the detailed capability table. That's not unusual for a model launch, but it's a real asymmetry worth naming every time it shows up: the parts of a release that are easy to quantify get a table, and the parts that are hard to quantify get adjectives.

Availability and pricing

Grok 4.6 is live today in Cursor, Grok Build, the API, and partners including OpenRouter, Vercel, and Cloudflare. Pricing holds at $2/$6 per million input/output tokens — unchanged from Grok 4.5's base rate — with a fast variant at a clean 2x ($4/$12). That's a simpler structure than 4.5's fast tier, which ran an asymmetric $4/$18 — a real pricing-structure change, not just a number update. A first-week promotion doubles included usage inside Grok Build and Cursor. No open-weights announcement accompanies this release — xAI/SpaceXAI's pledge, made in the same breath as joining the Open Secure AI Alliance, to eventually open Grok's weights remains unfulfilled for the flagship line as of this launch.

Update (August 13): ARC Prize verifies it, and the efficiency claim holds up

Independent confirmation arrived fast, if not on Harvey LAB. ARC Prize — the nonprofit that runs the ARC-AGI benchmark suite and independently verifies vendor submissions rather than accepting self-reports — published verified results for Grok 4.6: 87.5% on ARC-AGI-1 at $0.30/task, 67.1% on ARC-AGI-2 at $0.76/task, and 2.11% on ARC-AGI-3 at $5.6K total. This is a materially stronger source than xAI's own launch table: a "VERIFIED" badge rather than a vendor self-report, the same distinction this dataset has tracked all month between solid (verified) and hatched (self-reported) points on the ARC-AGI-3 leaderboard.

Worth the generational comparison, too: Grok 4.5's own verified scores were 85.7% / $0.33 (ARC-AGI-1), 52.6% / $0.78 (ARC-AGI-2), and 0.3% / $6.9K (ARC-AGI-3). Grok 4.6 improved every score while holding cost flat or lower — a real, independently confirmed generational gain, most dramatically on ARC-AGI-2 (52.6% → 67.1%) and, in relative terms, ARC-AGI-3 (0.3% → 2.11%, roughly 7x, though still a low absolute number).

The efficiency claim is real and specific: on ARC-AGI-3, Grok 4.6 with xhigh reasoning scored comparably to GPT-5.6 Sol with high reasoning, at $5.6K against Sol's $15.2K — roughly a third of the cost for a similar result. But read the chart before repeating that as a win. Every model clustered in that $1K–$20K cost band — Grok 4.6, GPT-5.5, GPT-5.6 Sol at both effort levels, Gemini 3.1 Pro — scores under 10% on ARC-AGI-3. The efficiency comparison is real within a cluster of models that are all still mostly failing the benchmark. Claude Opus 5 (High) sits alone near 30%, in a different part of the chart entirely, at a comparable order of magnitude in cost. "Efficient at 2%" and "capable at 30%" are different axes, and the verified chart makes that gap harder to miss than xAI's own table did.

What to expect next

  • Watch for independent verification on Harvey LAB specifically. ARC Prize covers ARC-AGI, not the legal-domain benchmark — a 6x gap over GPT-5.6 Sol there is still unconfirmed by anyone outside xAI. A real specialization or a benchmark GPT-5.6 Sol simply wasn't tuned for — third-party legal-AI evaluation would settle it either way.
  • Watch whether "third-party self-reported" stays the norm for the rest of this comparison table. ARC Prize closed that gap for ARC-AGI specifically. xAI's own footnote still concedes the other eight benchmarks are cited from rivals' published numbers, not run head-to-head — a methodology gap every lab in this comparison shares outside the one benchmark that now has independent verification.
  • Watch the self-verification claim get measured. "More self-testing on longer trajectories" is a real, checkable behavioral claim if someone builds an eval for it — right now it's an anecdote in a launch post.
  • Watch the open-weights pledge. xAI joined an alliance arguing for open, inspectable defensive AI while shipping its flagship closed for the second release running; the gap between that stated position and Grok's actual release pattern is worth tracking as its own story.

Read next