2026-08-14

Qwen3.8-27B: A 27B Model Beating a Model 30x Its Size on Most of Its Own Comparisons

AIOpen SourceModels🌍 Asia

Alibaba's Qwen team shipped open weights for Qwen3.8-27B — a dense, native multimodal model under Apache 2.0, with 262K native context extensible to roughly 1M tokens via YaRN. It's a small model by this month's standards, and that's exactly what makes its own benchmark tables worth reading closely: Alibaba is claiming a 27B model competes with, and mostly beats, closed frontier models dozens of times its size.

Counting wins in the coding, agent, and general table

Alibaba's own comparison table (Qwen3.8-27B vs. Qwen3.6-27B, Qwen3.7-Plus, Meta's Muse Glimmer-30B, and Claude Opus 4.6 Max) covers nine benchmarks across coding, agentic, and general reasoning. Counted directly: Qwen3.8-27B wins eight of nine outright — SWE-bench Pro (61.7), QwenSWEBench (79.0), CoWorkBench (70.7), JobBench (33.4), Agents' Last Exam (42.9), IFBench (79.5), and LiveCodeBench v6 (90.3), several by wide margins over its own predecessor Qwen3.6-27B. Opus 4.6 Max wins the other four: Terminal Bench 2.1 (78.2 vs. 73.0), NL2Repo-Bench (47.6 vs. 42.3), GPQA Diamond (91.3 vs. 89.2), and HLE (40.0 vs. 30.8). The pattern is specific, not blanket dominance: strong, often leading, performance on coding-agent and real-world office-work benchmarks. The gap to closed frontier models widens specifically on the hardest pure-reasoning evals — HLE most of all, where Opus 4.6 Max leads by nearly ten points. Still, a 27B model trading blows with, and mostly beating, a model reportedly 30x its size in its strongest categories is a genuinely notable efficiency result, regardless of where it trails.

The multimodal table, and where "outperforms Qwen3.7-Plus overall" gets thinner

On the 13-benchmark multimodal table, Qwen3.8-27B wins decisively on the agentic/embodied side — computer use (OSWorld-Verified, 84.3), browser use (WebArena, 64.8), mobile use (AndroidWorld, 81.9), application recreation (47.1), multimodal software engineering (SWE-MM, 38.6), visual web development (Vision2Web, 62.9), and the visual-math and general-visual-reasoning rows with chain-of-thought enabled. But Qwen3.7-Plus — a larger current sibling, not a superseded generation — wins document intelligence (OmniDocBench, 91.4), real-world perception (RealWorldQA, 86.9), embodied intelligence (ERQA, 69.8), and the ClawEval-MM average score. Alibaba's own launch copy calls Qwen3.8-27B a model that "outperforms Qwen3.7-Plus overall" — true as a composite claim, but worth reading against the actual row-by-row split. This isn't a clean sweep; it's a genuine capability trade, and "overall" is doing the same kind of compressing work we've flagged in other labs' headline framings this month.

What's still self-reported

Every number above comes from Alibaba's own comparison table — vendor-run, not independently verified, and the model card doesn't specify third-party reproduction of any of these results. That's worth stating plainly rather than repeating the numbers as settled fact, especially given how decisively Qwen3.8-27B claims to win most of its own comparisons. The weights being open and Apache 2.0-licensed at least makes independent reproduction possible now — a 27B dense model is small enough for outside labs and individual researchers to actually run and check.

What to expect next

  • Watch for independent benchmark reproduction. A 27B dense model is a realistic size to verify directly, which is exactly the test this specific set of claims needs.
  • Watch the HLE and GPQA gap specifically. Opus 4.6 Max's lead there suggests this Qwen generation is optimized hard for agentic/coding/office-work benchmarks at some cost to the hardest general-reasoning tests — worth tracking whether that's a deliberate trade or a genuine capability ceiling.
  • Watch adoption for lightweight, locally-deployable multimodal apps — Alibaba's own framing pitches this size specifically at builders shipping applications locally rather than through an API, and the real test of that pitch is what gets built with it.

Read next