FRONTIER ARCADE

270 AI RELEASES · 2022—2026
PLAY IT · CHART IT · SIZE IT · HARNESS IT · AUDIT IT
▶ TIMELINE CLASHbefore or after? 3 lives. the dates get closer…
A model appears. Did it release
BEFORE or AFTER the one above?

❤️❤️❤️ 3 lives · ⏱ 12s
streak = multiplier
every 10th round: BOSS — same lab!
HIGH SCORE: 0
0
– VS –
⚔ BOSS ×3 · SAME LAB ⚔
GAME OVER
📈 LABS × TIMEtap a quarter or drag to zoom · tap a dot for details
N.AMERICA EUROPE ASIA LANG IMG AUDIO WORLD
● lang ■ img ▲ audio ◆ world · ●filled = open-weight · ○hollow = proprietary
📏 MODEL × SIZE
big ● = total params · small ● + dashed line = active params (MoE only) · ●filled = open-weight · ○hollow = proprietary · color = region · shape = modality · log scale
🔧 HARNESS × SCORE
▮ solid = ARC PRIZE VERIFIED · ▯ hatched = self-reported by the harness author · ┆ gold = 95.4% human-expert baseline (Prime Intellect's chart)
Same benchmark, same public games. At launch every frontier model scored under 1%. None of these harnesses changed a single weight. Published human baselines vary by who ran the study and which testers they used — OpenAI separately estimates the average tester at 48% on this set. Treat every "vs. human" line here as specific to its own source, not a single agreed number.
NO SHARED SCOREharnesses this dataset tracks that never ran ARC-AGI-3 — different layer, same instinct
🗄 THE DATABASE
↑ CHART