Datapoint AI released what it's calling the largest open-source human-preference dataset for image generation: 2 million-plus pairwise annotations by real people, ranking 30 current image models across 10 use-case categories — marketing, product design, anime, and others — alongside a $1 million data grant for further work in the space. The dataset itself is on Hugging Face. The methodology is worth taking seriously on its own terms before getting to what the resulting rankings actually show.
Pairwise human preference is the right instrument, and it's rarer than it should be
Most image-model comparisons still lean on automated aesthetic scorers or CLIP-similarity metrics — proxies for what a person would actually prefer, not the preference itself. Datapoint's approach is closer to the LMSYS Chatbot Arena model for text: real people looking at two outputs side by side and picking one, aggregated into an Elo score across enough comparisons that noise averages out. At 2 million-plus annotations across 30 models, that's a substantial sample size for this kind of study. And the category breakdown (marketing, product design, anime, and seven others) means the ranking isn't collapsing every use case into one undifferentiated "quality" score the way a single aggregate number tends to.
Datapoint is also, worth naming plainly, a data-annotation company — publishing the largest available human-preference dataset for this task is a direct demonstration of the pipeline it sells, not a disinterested academic exercise. That doesn't make the data wrong, but it's the same kind of self-interest worth noting on any benchmark-builder's own claim about its benchmark, and readers should weigh the methodology on its merits rather than the "world's largest" framing alone.
The top seven are all closed
Counting directly down Datapoint's own Elo table: GPT Image 2 (high) leads at 1131, followed by Seedream 5.0 Pro (1127), Nano Banana 2 (1122), Reve 2.1 (1113), MAI-Image-2.5 (1109), Nano Banana Pro (1106), and GPT Image 1.5 (high) (1103). Every one of those seven is a closed, API-only model. The first genuinely open-weight model on the list — Hunyuanimage 3.0, from Tencent — sits at rank 17 (1066), roughly 65 Elo points back from the leader. FLUX.1 [schnell], Black Forest Labs' fast open-weight release, sits dead last at rank 30 (1000). That specific result is less a capability gap than an expected trade-off: "schnell" is explicitly optimized for speed over quality, and ranking last on a pure-preference benchmark is close to the intended shape of that trade rather than a surprise.
That said, the pattern is worth stating plainly rather than assuming: on this specific dataset, at this specific moment, human preference for general image-generation quality currently favors closed models across the board, with open weights occupying the middle and lower tiers rather than competing at the top. That's a different picture from language models, where open-weight releases have repeatedly matched or beaten closed frontier models on individual benchmarks this month. Whether image generation is a genuinely different competitive landscape, or simply earlier in the same catch-up curve text models have already run, remains an open question.
A narrow Elo band doesn't mean the field is close
The table's other widely-quoted feature is its compression: the whole 30-model spread runs from 1131 down to 1000 — a band of just 131 Elo points, from the leader to the deliberately-fast-and-cheap model at the bottom. It's tempting to read that as evidence the field is tightly bunched and no model is meaningfully ahead. That inference doesn't hold, though, because the width of an Elo table is largely a property of how the ratings were computed, not of how close the models actually are.
Elo scores have no absolute meaning, and their spread is not read off the data directly — it falls out of methodological choices: the initialization value, the K-factor (or the regularization strength in a Bradley–Terry fit), and the composition of the model pool. That's well established for rating systems generally. A tell here is that the last-place model, FLUX.1 [schnell], sits at exactly 1000 — the conventional default starting rating. That suggests the scale is anchored to that floor and the estimates pulled inward toward it, compressing the visible range rather than reporting a genuine 131-point ceiling on the differences.
The cleanest way to see this is to look at other human-preference arenas measuring the same thing. On LMArena's text-to-image board — roughly 5.9 million votes across 76 models — scores run from about 897 to 1382, a spread near 485 points. GPT Image 2 holds what that board describes as the single largest lead over second place in its history. Artificial Analysis's image arena tells the same story: its top model sits tens of Elo points clear of second and well over 100 clear of the models a few ranks down. Same underlying question — which image does a person prefer? — but very different spreads, because the rating machinery differs. So the comparison to chess Elo, where hundreds of points separate skill tiers, was never apples-to-apples: 131 points on Datapoint's scale and 485 on LMArena's are not the same units.
None of this makes Datapoint's ordering wrong. It means the narrow band is a fact about Datapoint's Elo computation, not a finding about the models — and "nobody's far ahead" is precisely the reading the raw point spread cannot support. On the larger public arenas, at least one model is quite clearly out in front.
The ranking is confirmed by other sources
The ordering here is confirmed by independent sources, so it doesn't need to be recomputed from Datapoint's annotations to be believed. Human-preference leaderboards for image models are published continuously, and at larger scale, by operators with no stake in Datapoint's data: LMArena's text-to-image board and Artificial Analysis's image arena run their own votes, and the broad picture Datapoint reports — closed models on top, open weights clustered lower — is the one they already show. What's distinctive about this release isn't the ordering but the open 2-million-vote dataset and the per-category breakdown behind it.
What to expect next
- Watch what the open dataset adds beyond the public arenas. The overall ordering already tracks what LMArena and Artificial Analysis independently report, so Datapoint's distinctive contribution is the open data and its per-category splits — worth watching whether those splits surface anything the aggregate leaderboards don't.
- Watch whether the closed-dominates-top pattern holds as open image models iterate. Hunyuanimage 3.0 at rank 17 is the current best open showing; whether the next generation of open image models closes that ~65-point gap to the leaders is a trackable, specific target.
- Watch what the $1M data grant actually funds. A grant announced alongside a benchmark release is a stated commitment with a use it can be held to — worth checking what gets funded and by when.