Datapoint AI released what it's calling the largest open-source human-preference dataset for image generation: 2 million-plus pairwise annotations by real people, ranking 30 current image models across 10 use-case categories — marketing, product design, anime, and others — alongside a $1 million data grant for further work in the space. The dataset itself is on Hugging Face. The methodology is worth taking seriously on its own terms before getting to what the resulting rankings actually show.
Pairwise human preference is the right instrument, and it's rarer than it should be
Most image-model comparisons still lean on automated aesthetic scorers or CLIP-similarity metrics — proxies for what a person would actually prefer, not the preference itself. Datapoint's approach is closer to the LMSYS Chatbot Arena model for text: real people looking at two outputs side by side and picking one, aggregated into an Elo score across enough comparisons that noise averages out. At 2 million-plus annotations across 30 models, that's a substantial sample size for this kind of study, and the category breakdown (marketing, product design, anime, and seven others) means the ranking isn't collapsing every use case into one undifferentiated "quality" score the way a single aggregate number tends to.
Datapoint is also, worth naming plainly, a data-annotation company — publishing the largest available human-preference dataset for this task is a direct demonstration of the pipeline it sells, not a disinterested academic exercise. That doesn't make the data wrong, but it's the same kind of self-interest worth noting on any benchmark-builder's own claim about its benchmark, and readers should weigh the methodology on its merits rather than the "world's largest" framing alone.
The top seven are all closed
Counting directly down Datapoint's own Elo table: GPT Image 2 (high) leads at 1131, followed by Seedream 5.0 Pro (1127), Nano Banana 2 (1122), Reve 2.1 (1113), MAI-Image-2.5 (1109), Nano Banana Pro (1106), and GPT Image 1.5 (high) (1103). Every one of those seven is a closed, API-only model. The first genuinely open-weight model on the list — Hunyuanimage 3.0, from Tencent — sits at rank 17 (1066), roughly 65 Elo points back from the leader. FLUX.1 [schnell], Black Forest Labs' fast open-weight release, sits dead last at rank 30 (1000) — though that specific result is less a capability gap than an expected trade-off: "schnell" is explicitly optimized for speed over quality, and ranking last on a pure-preference benchmark is close to the intended shape of that trade rather than a surprise.
That said, the pattern is worth stating plainly rather than assuming: on this specific dataset, at this specific moment, human preference for general image-generation quality currently favors closed models across the board, with open weights occupying the middle and lower tiers rather than competing at the top. That's a different picture from language models, where open-weight releases have repeatedly matched or beaten closed frontier models on individual benchmarks this month — worth tracking whether image generation is a genuinely different competitive landscape or whether it's simply earlier in the same catch-up curve text models have already run.
The whole field is closer together than "SOTA" implies
The second finding is arguably more interesting than the ranking order itself: the entire 30-model spread runs from 1131 to 1000 — a band of 131 Elo points across every model Datapoint tested, from the acknowledged leader to the deliberately-fast-and-cheap model at the bottom. For comparison, a meaningful skill gap in a game like chess Elo typically spans many hundreds to over a thousand points; a 131-point spread across an entire field, top to bottom, describes a set of models that are, in human-preference terms, much closer together than "state of the art" headlines about any single model tend to suggest. Rank 15 (FLUX.2 [max], 1072) and rank 20 (Z-Image Turbo, 1057) differ by less than the gap between ranks 1 and 2 — which is a useful check against reading any specific rank ordering in the middle of the table as a meaningful capability difference rather than noise within a tightly clustered field.
What's not yet independently checked
This is Datapoint's own dataset, own methodology, and own published Elo table — the annotations themselves being open and downloadable is what makes independent reproduction possible, but nobody outside Datapoint has run that reproduction yet as of this post. Worth applying the same standard here as to any self-published ranking: the open data is real credibility groundwork, not a substitute for someone else recomputing the Elo scores from the raw annotations and confirming the ordering holds.
What to expect next
- Watch for independent Elo recomputation from the raw annotation data. The dataset is open specifically to make this possible — whether the published ranking survives someone else's independent pass is the real test of the release.
- Watch whether the closed-dominates-top pattern holds as open image models iterate. Hunyuanimage 3.0 at rank 17 is the current best open showing; whether the next generation of open image models closes that ~65-point gap to the leaders is a trackable, specific target.
- Watch what the $1M data grant actually funds. A grant announced alongside a benchmark release is a stated commitment with a use it can be held to — worth checking what gets funded and by when.
References: Datapoint AI on X — announcement · Hugging Face — datapointai/text-2-image-human-preferences-2m · related coverage: FLUX 3 · DiffusionGemma · Qwen3.8-27B · Frontier Arcade: trends & predictions