2026-08-01

Qwen3.8-Max Showed Its Work — and Led Almost Every Multimodal Chart Doing It

AIBenchmarks🌍 Asia

Two weeks ago, Alibaba previewed Qwen3.8-Max at WAIC Shanghai with a 2.4-trillion-parameter figure and a claim to be "second only to Fable 5" — and nothing to check it against: no benchmark table, no model card, no active-parameter count, no license. It was, fairly, criticized as a claim with nothing behind it. This week the benchmarks arrived, alongside a launch tweet confirming pricing and a genuinely notable openness commitment: open weights for both Qwen3.8-Max and a new, smaller Qwen3.8-27B, promised next week.

What the tables actually show

Across two large comparison charts against Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Pro, and GPT-5.6 Sol, Qwen3.8-Max doesn't win everything — but where it wins, it wins by a real margin, and the pattern is specific rather than diffuse.

Where it leads outright: PaperBench (research reproduction) at 93.0, ahead of GPT-5.6 Sol's 90.5 and Fable 5's 88.8; IFBench (instruction following) at 82.8, clear of Sol's 72.7; HealthBench at 60.2 against Sol's 55.3; PLawBench and PRBench-Finance, both led by Qwen. And the multimodal chart is close to a clean sweep: MMMU-Pro, MathVision, BabyVision, HiPhO, PhyX, SLAKE, OSWorld-Verified, Parametric CAD Bench, VideoMME, VideoMMMU, MMVU, MLVU — Qwen3.8-Max tops all of them, ahead of Gemini 3.1 Pro and GPT-5.6 Sol specifically, not just older models.

Where it doesn't: SWE-bench Pro sits at 67.7 against Fable 5's 80.0 — a 12-point gap that's hard to read as competitive. FrontierSWE is 73.5 against Fable 5's 88.8. HLE, the hardest general-reasoning benchmark on the sheet, is Qwen's weakest relative showing at 43.6 — lowest of the five models compared, behind even Opus 4.8's 45.7. AndroidWorld and ScreenSpot Pro both go to Fable 5 as well.

Put together, that's a specific and checkable claim, not the vague ranking line from two weeks ago: Qwen3.8-Max is a genuine multimodal and professional-domain leader, and a clearly second-tier coding agent behind Fable 5 specifically. Worth remembering these are still Alibaba's own internal evals, not an independent leaderboard like Artificial Analysis or LMArena — the July 19 preview drew criticism for exactly that gap, and shipping a comprehensive table doesn't itself resolve who ran the eval.

The pricing puts real weight behind the multimodal claim

Qwen3.8-Max launched at $2.00 input / $6.00 output per million tokens, with implicit caching at $0.25. Lay that against the pricing table this blog put together a couple of days ago: that's identical to Grok 4.5, cheaper than GPT-5.6 Terra ($2/$12) on output, and a fraction of GPT-5.6 Sol ($5/$30) or Claude Fable 5 ($10/$50) — the two models Qwen3.8-Max is actually beating on several of the multimodal and professional-domain benchmarks above. A model priced at mid-tier that leads flagship-tier models on a specific, real cluster of benchmarks is a much more interesting story than "second only to Fable 5," and it's the one the numbers actually support.

Open weights, twice — if the date holds

The launch tweet commits to open-weighting both the full 2.4T Qwen3.8-Max and a new, considerably smaller Qwen3.8-27B, "next week." That's a meaningfully bigger openness bet than most labs make at once — releasing a flagship-scale model's weights alongside a deployable small model in the same window gives adopters a genuine choice of scale rather than one size-fits-all download. It's also a promise this exact model already broke once: the July 19 preview said open weights were "planned for official release" with no date, and slipped past that framing for two weeks before this benchmark drop. Worth treating "next week" as a date to watch, not one to bank on.

The demonstrations are not the benchmarks

The launch also leans on capability anecdotes that don't appear in any table above: "10+ days of self-evolving development, from empty folder to production without hand-holding," 500+ turns of autonomous chip-design optimization, 365 days of simulated e-commerce strategy. These are real and worth watching — Alibaba published a GitHub trace for the coding claim rather than just asserting it — but they're demonstrations of a single run, not a scored, repeatable benchmark the way PaperBench or SWE-bench Pro are. Keep them in a separate mental bucket from the numbers above until someone outside Alibaba reproduces one.

What to expect next

  • The "second only to Fable 5" framing quietly disappears from Qwen's own marketing. It didn't survive contact with a benchmark table that shows a much more specific, two-sided story — multimodal leader, second-tier coding agent — and that's a harder line to put on a slide than a single ranking claim.
  • Watch the actual open-weights date. This model already missed one implicit deadline; whether Qwen3.8-Max and Qwen3.8-27B ship open weights next week, as promised, is the test of whether this launch's other numbers deserve the benefit of the doubt.
  • Independent verification of the multimodal sweep matters more than any other number here. If Artificial Analysis or a similar third party reproduces even half of the vision/video leads once weights are public, this becomes the most credible multimodal open-weight claim of the year; if it doesn't reproduce, "own internal evals" was doing more work in this table than the chart lets on.

References: Alibaba Qwen (@Alibaba_Qwen) — Qwen3.8-Max launch announcement · Qwen — official blog post · MarkTechPost — the July 19 preview and "zero benchmarks" context · related coverage on this blog: The price war nobody is actually fighting