xAI shipped Grok 4.6 today, a day after Grok Bot's public beta launch — the same shipping cadence we flagged as the point of the last four months of SpaceXAI/Cursor dealmaking. The positioning is explicit: "a particular focus on long-running agents and more ambitious interactive and visual work," staying with complex multi-step tasks — research, codebase work, or turning an idea into a finished artifact. xAI's own headline claim is that Grok 4.6 "matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index." That's true. It's also the least interesting sentence in a nine-benchmark table that tells a more specific story than the composite score does.
What the table actually shows
Reproducing xAI's own comparison table (Grok 4.6 High vs. Grok 4.5 High vs. GPT-5.6 Sol Max vs. Claude Fable 5 Max, third-party figures per xAI "the best of self-reported or publicly available results" — worth flagging that caveat up front, since it means xAI chose which of its rivals' own numbers to cite, not that xAI ran the comparisons itself):
| Benchmark | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Count the wins and the composite-score framing gets less flattering: Fable 5 Max leads on five of the nine metrics (AA Intelligence Index, CursorBench, FrontierCode, APEX-Agents, APEX-SWE), GPT-5.6 Sol Max leads on two (DeepSWE, Terminal-Bench), and Grok 4.6 leads on three (GDPVal-AA, AA-Briefcase, Harvey LAB). "Matches GPT-5.6 Sol" is accurate on the one composite metric xAI chose to headline; it's not the metric where Grok 4.6 actually looks strongest, and it's not close to describing Fable 5's position, which leads more often than either of the models xAI compared itself to directly.
The real pattern: knowledge work and one specific vertical, not classic coding agents
The shape of where Grok 4.6 actually wins is more informative than the headline. Its three clean leads — GDPVal-AA (a real-world economic-value task index), AA-Briefcase (knowledge-work tasks), and Harvey LAB (Vals) — cluster around broad office and professional work, with a standout on Harvey LAB specifically: a legal-domain benchmark where Grok 4.6 posts 15.8% against GPT-5.6 Sol's startlingly low 2.5%. That's a 6x gap on one specific vertical, which reads either as a genuine specialization or as a benchmark GPT-5.6 Sol simply wasn't tuned for — worth independent scrutiny either way.
Where it doesn't lead is the classic coding-agent cluster: DeepSWE, Terminal-Bench, CursorBench, FrontierCode, and APEX-SWE all go to either Fable 5 or GPT-5.6 Sol. That sits in some tension with the announcement's own framing — "working across a codebase" is explicitly named as a target use case — while the numbers say Grok 4.6 is a competent-but-not-leading coding agent and a genuinely strong knowledge-work and legal-domain model. The interesting Grok 4.6 story is narrower and more specific than "frontier intelligence" implies, and also more useful if it holds up: a model that's good at office and professional tasks isn't a lesser story than a coding benchmark leader, it's a different one.
Training: the model teaching its own successor's data
xAI describes a longer supplemental training run than 4.5's, with curated model-generated data for reasoning and technical concepts, plus an improved optimizer and recipe as the SFT/RL foundation. The more interesting detail: Grok 4.5 was used to regenerate the SFT trajectories for 4.6 — across reasoning efforts, agent harnesses, and domains — with model-based filtering to drop bad traces. That's the same self-bootstrapping pattern we've now seen from Meta (Spark 1.1 grading Spark 1.2's candidate solutions) and implicitly from Liquid's MOPD: use the previous generation as a teacher for the next generation's training data, rather than only generating from scratch. It's becoming a standard move across labs, not a one-off technique — worth tracking as its own trend line.
RL training covered "a wide range of agentic RL tasks" — knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and CAD. xAI also reports Grok 4.6 showing more self-testing and verification behavior on longer trajectories — "checking its own work before moving on" — which is a qualitative, unbenchmarked claim, not a scored metric, and worth treating that way until someone measures it directly.
The safety section says a lot by saying little
Compare the specificity gap: nine quantified capability benchmarks with exact scores, against a safety paragraph that says safeguards were "improved and calibrated in line with the model's capabilities," backed by "our widest-ever suite of pre-deployment testing" and "extensive post-deployment and third-party testing" — no numbers, no named third parties, no published red-team results alongside the detailed capability table. That's not unusual for a model launch, but it's a real asymmetry worth naming every time it shows up: the parts of a release that are easy to quantify get a table, and the parts that are hard to quantify get adjectives.
Availability and pricing
Grok 4.6 is live today in Cursor, Grok Build, the API, and partners including OpenRouter, Vercel, and Cloudflare. Pricing holds at $2/$6 per million input/output tokens — unchanged from Grok 4.5's base rate — with a fast variant at a clean 2x ($4/$12). That's a simpler structure than 4.5's fast tier, which ran an asymmetric $4/$18 — a real pricing-structure change, not just a number update. A first-week promotion doubles included usage inside Grok Build and Cursor. No open-weights announcement accompanies this release — xAI/SpaceXAI's pledge, made in the same breath as joining the Open Secure AI Alliance, to eventually open Grok's weights remains unfulfilled for the flagship line as of this launch.
What to expect next
- Watch for independent verification, especially on Harvey LAB. A 6x gap over GPT-5.6 Sol on one legal benchmark is either a real specialization worth confirming or a sign the comparison set favors Grok 4.6's training data in a way that won't generalize — third-party legal-AI evaluation would settle it either way.
- Watch whether "third-party self-reported" stays the norm for these comparison tables. xAI's own footnote concedes it's citing rivals' published numbers, not running head-to-head tests itself — a methodology gap every lab in this comparison shares, and one that keeps these charts a step short of independent benchmarking.
- Watch the self-verification claim get measured. "More self-testing on longer trajectories" is a real, checkable behavioral claim if someone builds an eval for it — right now it's an anecdote in a launch post.
- Watch the open-weights pledge. xAI joined an alliance arguing for open, inspectable defensive AI while shipping its flagship closed for the second release running; the gap between that stated position and Grok's actual release pattern is worth tracking as its own story.
References: xAI — Introducing Grok 4.6 · SpaceXAI on X — benchmark table · related coverage: Grok Bot: The Product That Was Always the Point of the Cursor Deal · The Open Secure AI Alliance · Meta Muse Glimmer · Frontier Arcade: trends & predictions