▮ solid = ARC PRIZE VERIFIED · ▯ hatched = self-reported by the harness author · ┆ gold = 95.4% human-expert baseline (Prime Intellect's chart)
Same benchmark, same public games. At launch every frontier model scored under 1%. None of these harnesses changed a single weight. Published human baselines vary by who ran the study and which testers they used — OpenAI separately estimates the average tester at 48% on this set. Treat every "vs. human" line here as specific to its own source, not a single agreed number.
NO SHARED SCOREharnesses this dataset tracks that never ran ARC-AGI-3 — different layer, same instinct