2026-09-12

This Design Benchmark Tested Six Open-Weight Models and Seven Closed Ones. Only Two Were Worth Their Price — One of Each

AIModelsBenchmarks🌍 Global

The chart is unusually blunt for a benchmark. Most leaderboards rank models top to bottom and let the reader supply their own judgment about whether the gap between rank three and rank eight is worth paying for. OpenDesign Arena instead plots score against price on a log-scaled x-axis, draws a dashed line at the score of its cheapest strong performer, and shades everything below and to the right of it in pink. The shading has one meaning: every model in that zone scored the same as or worse than a model that cost dramatically less. Of the 13 models tested, 11 sit inside the pink.

OpenDesign Arena's score-versus-price chart (French-language version of the site): GPT-6 Astra sits alone at the top right with the highest average score at the highest price; DeepSeek V4.1 Flash sits alone at the bottom left, circled as the best-value pick, with a near-identical score at a small fraction of the price. A dashed line runs from DeepSeek V4.1 Flash's score across the chart, and a shaded pink region below it contains the other eleven models, including Claude Fable 5.1, GPT-5.6 Sol, Grok 4.6, Qwen 3.8-Max, Hunyuan H4 Preview, Gemini 3.8 Flash, GLM-5.3 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, Muse Spark 1.3, and Kimi K3 — all scoring the same or worse than DeepSeek V4.1 Flash while costing more.

The natural reading of "the cheapest model nearly ties the most expensive one" is an open-versus-closed story: open-weight models are catching up, closed frontier labs are charging a premium they can no longer fully justify. That story is half right. Sorting all 13 models by license, rather than by score or price, is what actually explains the chart — and it isn't the story either side would expect.

Thirteen models, sorted by who can see the weights

Six of the thirteen models here are open-weight. Seven are closed, API-only products.

Open-weight: DeepSeek V4.1 Flash and DeepSeek V4-Flash (both MIT-licensed, on Hugging Face), DeepSeek V4-Pro (MIT — "the least encumbered choice available," this blog noted at its launch), GLM-5.3 Flash (MIT-licensed, notably not the Apache 2.0 license Z.ai used for its other recent releases), Hunyuan H4 Preview (Tencent's newest, Apache 2.0, released to Hugging Face on August 28), and Kimi K3 (Moonshot's own description: the largest open-weight model yet).

Closed: GPT-6 Astra and GPT-5.6 Sol (OpenAI, API-only), Claude Fable 5.1 (Anthropic, API-only), Grok 4.6 (xAI shipped this one with, in this blog's own words at launch, "no open-weights announcement"), Gemini 3.8 Flash (Google, whose open-weight line is branded separately as Gemma rather than Gemini), Qwen 3.8-Max (Alibaba's API-only flagship tier — a separate, text-only open checkpoint of the underlying model exists, but it lagged the closed API version by two weeks and isn't the same multimodal product being benchmarked here), and Muse Spark 1.3 (Meta's model, which shipped closed on September 3 after already promising, and not delivering, open weights for its predecessor Muse Glimmer three weeks earlier — the same "open weights... coming soon" language recycled from one broken timeline to the next).

Now overlay that split onto the chart's two undominated points. GPT-6 Astra, the highest scorer at the highest price, is closed. DeepSeek V4.1 Flash, the cheapest model that still scores near the top, is open. If the story stopped there, it would be a clean parable: pay a premium for a closed frontier model, or get almost the same result for nearly free from an open one.

It doesn't stop there, because of what happens to the other eleven.

Open-weight buys you nothing by itself

Five of the six open-weight models tested — Hunyuan H4 Preview, DeepSeek V4-Pro, DeepSeek V4-Flash, GLM-5.3 Flash, and Kimi K3 — sit inside the shaded, dominated zone along with the closed models. Being open-weight didn't put any of them on the frontier. Two of those five are DeepSeek's own prior releases: the company that produced the chart's cheap frontier point also produced two of the chart's dominated points, using the same MIT license, the same lab, the same general architecture family. DeepSeek V4.1 Flash isn't cheap and capable because DeepSeek open-weights its models — V4-Pro and V4-Flash are open-weight too, and neither escapes the pink. It's cheap and capable because it's DeepSeek's newest architecture specifically, built around an encoder-decoder design this blog covered in detail from DeepSeek's own model card, which happens to also be open-weight. The licensing is incidental to the result; the architecture generation is what did the work.

That matters because "open-weight" is often treated as a proxy for "cheap and efficient" in industry commentary, and this chart is a clean counterexample: five out of six open models tested cost more and scored worse than the sixth. For DeepSeek's own two dominated entries, that's a newer model beating its own predecessors. For Hunyuan H4 Preview, GLM-5.3 Flash, and Kimi K3, it's something plainer: being open-weight didn't buy them competitiveness against a rival lab's model at all, open or closed. Licensing tells you nothing about which side of that line a given open model falls on.

Closed doesn't buy you the frontier either

The mirror case is just as clean. Five of the seven closed models — GPT-5.6 Sol, Grok 4.6, Qwen 3.8-Max, Gemini 3.8 Flash, and Muse Spark 1.3 — are also inside the shaded zone. Paying for API-only access to a closed model bought none of them a spot outside the dominated region either. Only GPT-6 Astra, among seven closed products, earned its price on this specific benchmark.

Muse Spark 1.3 is the sharpest version of this. It's closed, and Meta has now twice told its own users that open weights are coming without giving a license or a date — first for Muse Glimmer in August, then again for Spark 1.3 three weeks later, reusing nearly identical language both times. A model that's closed today, with an unfulfilled promise to open it, and that scores in the bottom half of this chart at a price higher than the open-weight frontier point, gets none of the usual advantages either side of the open/closed argument claims: not the transparency and low cost open advocates point to, and not the performance premium that's supposed to justify staying closed.

What actually separates the two dots from the other eleven

None of the eleven dominated models are dominated because of their license. They're dominated because, on this specific benchmark — building web apps, dashboards, and landing pages, scored on a 100-point rubric split between requirement fulfillment and design quality — a newer architecture from one lab (DeepSeek's V4.1 Flash) and a larger, more expensive one from another (OpenAI's GPT-6 Astra) both beat everything else tested, including their own makers' other products and everyone else's closed and open releases alike. DeepSeek V4.1 Flash reaches 98% of GPT-6 Astra's score at 1.4% of its price — 81.2 against 82.7, at 0.023perfinisheddesignagainst0.023 per finished design against 1.61 — while Claude Fable 5.1, a closed model priced at $3.66 per design, scores lower than the open one at more than 150 times the cost. That's the actual shape of a Pareto-style plot: a model earns a place outside the dominated zone by beating the frontier on at least one axis, and the eleven models in between — five open, six closed by count once Astra and V4.1 Flash are set aside — simply don't, regardless of who can see their weights.

What was actually tested, and by whom

OpenDesign Arena scores models on ordinary front-end design work: web apps, dashboards, mobile screens, landing pages, built from a task brief. Each finished artifact is scored out of 100. The first check is mechanical and happens before anything is graded — the harness loads the generated code in a headless Chromium instance, and a blank viewport, a broken import, an unrendered markdown block, or an uncaught JavaScript exception produces an automatic zero. Only code that survives this gate gets scored at all.

Of the 100 points available to a surviving artifact, 30 go to requirement fulfillment — whether the required pages, content, states, and interactions are present — and 70 go to design quality: layout, hierarchy, style fit, color and contrast, image relevance. That second category is where a benchmark like this usually gets vague, and OpenDesign's answer is an LLM judge validated against human raters rather than left to grade itself unchecked. In a separate reliability study, ten human evaluators (three professors and seven graduate students) and a GPT judge each did pairwise comparisons of the same outputs. The two humans agreed with each other 68.7% of the time; the GPT judge agreed with the human panel 80.9% of the time — a higher rate than the humans managed with each other. That says the judge is at least as consistent as the panel checking it, not that either is measuring some ground truth beyond dispute, but it's a real, checkable validation rather than an unsupported claim.

Independent coverage of this leaderboard update converges tightly on three data points — the bar this blog treats as trustworthy for a third-party number it can't pull from the primary source directly, since OpenDesign's own site didn't load for this post and the figures below come from multiple outlets citing the same table:

  • GPT-6 Astra: 82.7 average (26.5/30 requirement fulfillment, 56.2/70 design quality), $1.61 per design, 11.1 minutes.
  • DeepSeek V4.1 Flash: 81.2 average, $0.023 per design, 5.3 minutes.
  • Claude Fable 5.1: 80.3 average, $3.66 per design, 12.8 minutes.

This post won't assign the other ten models precise scores it couldn't verify against OpenDesign's own table, but their rank order and rough clustering are legible directly from the chart, and the open/closed licensing status above is drawn from this blog's own prior reporting and each maker's published license, not from OpenDesign.

The benchmark itself isn't neutral on this question

It's worth being as skeptical of OpenDesign as this blog would be of a lab's own chart, for a different reason than usual. OpenDesign is the open-source, Apache-licensed product of nexu-io: a local-first design tool, pitched as a Claude Design alternative, that turns a coding agent into a design engine and works BYOK (bring your own key) across more than 20 different model providers. A company whose product's entire value proposition is "plug in whichever model you like, we don't care which" has a real commercial interest in a headline that reads "the cheap open option is 98% as good for 1.4% of the price" — that's the exact pitch for a model-agnostic tool, independent of whether it's true. It's not that the incentive necessarily distorts the automated, mechanical part of the score, and the human-validation study is a genuine, checkable methodology rather than a marketing claim standing alone. But a benchmark sponsor's business model is data, the same way a lab's own launch chart is data, and a BYOK design-tool vendor benefits more from "openness wins on cost" being believed broadly than the actual, messier finding here — that licensing predicts almost nothing, and only two specific models out of thirteen, one open and one closed, are worth choosing at all.