xAI launched Grok 4.7 with a specific, checkable claim: "a notable improvement over Grok 4.6 at the same price and speed." Its own seven-benchmark table backs that up completely — Grok 4.7 xHigh beats Grok 4.6 High on every single row, from CursorBench (46.3% vs 40.4%) to Terminal-Bench 4.0 (38.0% vs 20.3%), at the identical 6 per-million-token price. That part of the claim is simply true.
The table also includes GPT-5.6 Sol Max (20) and Claude Fable 5.1 Max (50), and the tally there is the more useful number. Against Sol, Grok 4.7 wins 5 of 7 rows — losing only DeepSWE v1.1 and HealthBench Pro — at half the price. Against Fable 5.1, it's 3 wins to 4, at a fifth of the price: Grok 4.7 leads on DeepSWE, EEBench, and Harvey Legal Agent, and trails on CursorBench, AA-Briefcase, Terminal-Bench, and HealthBench — the Terminal-Bench gap is the largest in the table, 38.0% against 57.9%. That's a genuinely strong price-to-performance position, not just a marketing line — xAI just didn't lead with it.
One footnote is worth a second look. The DeepSWE v1.1 row carries an asterisk reading "High Effort" next to Grok 4.7's 71.0% — meaning that specific number was run at the "High" tier, not the "xHigh" tier every column header, and every other row, says Grok 4.7 was tested at. No reason is given for the swap, and it's the one row where Grok 4.7 loses to Sol (72.7%). Whether xHigh would have closed that gap or widened it isn't something the table lets a reader check.
This blog covered Grok 4.6's launch making a similarly accurate but selectively-framed claim — "matches GPT-5.6 Sol" was true and also the least interesting line in that table. Same pattern here: the headline claim holds, and the more informative number is one row down.