2026-08-10

What Four Years of Model Releases Say About Where AI Is Going

AIData🌍 Global

The arcade is a curated dataset of 240 model releases from February 2022 to August 2026 (247 entries counting the 7 industry events), each tagged with lab, continent, modality, and open/closed access. Is that enough data to stop guessing and start fitting curves? I decided to find out… with Fable 5. Only time will tell us if the predictions are right.

The release rate is exponential (and July broke the curve)

Fitting an exponential to releases per quarter gives +17% growth per quarter — a doubling roughly every four to five quarters — with R²=0.86 on the log scale. That alone would be remarkable. But the curve keeps being outrun: July 2026 brought 39 releases — more than one a day — and Q3 hit 48 releases by August 10, a pace of ~107 per quarter with seven weeks still to run, against a fitted prediction of 33.

Bar chart of AI releases per quarter from 2022 to 2026, with Q3 2026 already at 48 by August 10 and an exponential fit line extending as a forecast into 2027

One honest caveat: some of this is a curation artifact. The dataset logs 2026 at finer granularity than 2022 — variants, previews, and pricing tiers all count now in a way that "GPT-4 shipped" didn't (and last week's audit backfilled a dozen early open-weight releases, which is exactly why the 2022 baseline moved since the first version of this post). The exact doubling time is soft. What's not soft is the next chart.

A day between releases

The median gap between any two consecutive industry releases went from 23 days in 2022 to 1 day in 2026 — a 23× compression. (The 2022 figure moved from 25 in the first version of this post: the audit backfills added early releases, which is what re-measuring your own baseline costs.) Note the log scale: on a linear axis the 2023–2026 bars would be indistinguishable slivers.

Bar chart on a log scale showing the median days between consecutive industry releases falling from 23 in 2022 to 1 in 2026

And the timing isn't random. Scanning flagship releases for rival launches within 48 hours turns up fourteen pairs since mid-2025, trending toward same-day: GPT-5.3-Codex landed the same day as Claude Opus 4.6, DeepSeek v4 shipped the day after GPT-5.5, GLM-5.2 launched the day after the Fable 5 suspension, Qwen3.8-Max-Preview arrived two days after Kimi K3 — and this week FLUX 3 dropped two days after Qwen Image 3.0. Release dates have become counterprogramming.

Asia takes the #2 spot — through open weights

Asia went from a single release in 2022 (GLM-130B) to 31% of everything shipped in 2026, overtaking Europe in 2025 and never looking back.

Stacked bar chart of release share per continent per year, showing Asia growing from 11 percent in 2022 to 31 percent in 2026

The more interesting split is how each region ships. In 2026, only 27% of North American releases are open-weight, versus 56% for Asia and 77% for Europe. The overall open share actually fell (73% in 2024 down to 42% in 2026) as the US frontier closed up — but the biggest open models are now frontier-scale: Kimi K3 at 2.8T parameters is the largest open-weight model ever, and Chinese open models reportedly account for ~61% of OpenRouter tokens. Open didn't lose share of quality; it lost share of count while gaining the top of the leaderboard.

The specialist wave

The most recent shift: models built for one job. Counting code, cybersecurity, OCR, formal math, voice, world-model, and agent-specialist releases per half-year:

Bar chart of specialist model releases per half-year, rising to 12 in the first half of 2026 and 13 in the second half by August 10

At most two per half-year through 2025, then a jump to 12 in 2026H1 — and 2026H2 is already at 13 by August 10 (9 in July, 4 more in the first ten days of August: a safety classifier, a coding agent's co-trained model, and two formal-logic specialists). The categories are multiplying too: to code, cyber, OCR, math and world models, the past six weeks added content-safety classification (Shieldstral) and formal deductive reasoning (TwiL-LM). They also cluster suspiciously: two OCR models within two days (Baidu Unlimited-OCR, Mistral OCR 4), and two cybersecurity models on the same day (Fugu-Cyber and Gemini 3.5 Flash Cyber, July 21). Every 2026 cyber model is gated behind trusted-partner access — specialization and access control are arriving together.

Predictions

Two already came true — in the same week

The first version of this post flagged an anomaly: Black Forest Labs had gone silent for six months, when its own release cadence predicted something around March. The bet, as written then: "either FLUX.3 is big, or something corporate is brewing." Two days later, on July 23, FLUX 3 landed — and it settled on big: not an incremental image model but a multimodal "visual intelligence" pivot spanning video, audio, and robot action prediction, with a variant already running on Audi production lines.

Then it happened again. This post flagged Anthropic as overdue — projected July 12 by its own cadence — and named its overshoot "the one to watch" if the BFL pattern held. On July 24, twelve days past projection, Claude Opus 5 landed: not an incremental bump but a model that debuted #1 on the Artificial Analysis index, above Anthropic's own Mythos-class Fable 5, at half the price. Late, and big — on pattern, twice in one week.

Worth being precise about what this validates. Cadence math can't tell you what a lab will ship — nothing in the gap statistics predicted robots, or a flagship-beating mid-tier. What it can do is flag that something is due: a lab with an established rhythm going quiet past its usual interval is either winding down or winding up, and for both BFL and Anthropic, "winding down" was implausible. That's the actual pattern worth keeping: silence from a lab with an established cadence is itself a signal, and the longer the silence runs past the projected date, the bigger the eventual release tends to be — they didn't skip a release, they saved it up. Next test of the rule: Google, projected mid-August.

The math I'd actually bet on

  • Release cadence per lab. Taking each lab's median of its last five inter-release intervals: Google projects to mid-August, Z.ai to mid-August, xAI to late September, DeepSeek to November — and Alibaba's median gap is down to two days, which is less a projection than a standing appointment. (Anthropic's entry resolved July 24 — see above.)
  • Q3 lands north of 90. The fit says 33, but July alone already brought 38; at that pace the quarter clears 100, though saturation may pull it back. Once the median gap touches ~1 day, discrete releases stop making sense and labs shift to continuously versioned channels — the GPT-5.x monthly increments are already halfway there.
  • Open weights rebound toward ~50% by 2027. Kimi K3's weights land July 27, Qwen3.8's open release is promised, Europe ships 75% open, and the quarterly open share already recovered from 33% (Q1) to 50% (Q3 so far).
  • Cyber becomes a permanently gated category. Three for three gated in 2026, plus the post-suspension severity framework the big labs are co-proposing. Expect "government and trusted partners only" to become a standard release tier, not an exception.
  • World models break out in 2027. Cosmos 3 shipped with 20+ industrial partners, HY-World 2.0 targets games, and Qwen-AgentWorld's result — simulated RL beating real-environment RL on some benchmarks — is the kind of enabling result that precedes a wave.
  • Orchestrators over monoliths. Fugu Ultra beats Opus 4.8 on SWE-Bench Pro by composing other models; Kimi K2.6 runs 300-agent swarms; Muse Spark is trained to orchestrate subagents. The unit of capability is shifting from "a model" to "a system that routes models" — which, incidentally, makes datasets like this one harder to curate every month.

Update (Aug 1, 2026): scoring the board

Nine days on, here's where the board stands — the two dated bets above resolved as hits, and the trend calls are already collecting confirming releases:

  • Cyber-gating held. MAI-Cyber-1-Flash (Microsoft) shipped July 27 — closed, like the three before it. Four cyber models, four gated. On thesis.
  • Orchestrators kept coming. Sakana's Fugu-Ultra v1.1 landed July 24 — another turn of the trained-orchestrator crank.
  • The open rebound is tracking. Q3's open-weight share is already at 50%, and six of the eleven releases since July 23 are open (LLaDA2.2-flash, Instella-MoE, A.X K2, Inkling-Small, DeepSeek-V4-Flash, and MiniMax-H3).
  • Counterprogramming reached the video labs — with the openness axis built in. On July 31, MiniMax shipped MiniMax-H3 (Hailuo 3.0) — an omni-modal video model (native stereo audio, up to 2K) with open weights landing August 3, at roughly 3× under the going rate — the same day ByteDance launched its closed Seedance 2.5. Same product category, same day, opposite bets on openness. The rival-launch pattern now sorts labs by strategy, not just timing.
  • One call is visibly stressed — and it proves the meta-rule. Alibaba's "two-day standing appointment" stalled: no release since July 21, eleven days of silence against a ~1-day cadence. By this post's own logic that isn't a refutation — silence from a lab with an established cadence is itself a signal, and Alibaba going quiet reads as winding up, not down. Watch that space.
  • DeepSeek jumped its own queue. Projected to November, it instead shipped DeepSeek-V4-Flash on July 31 — a Flash variant rather than a new flagship, so the November flagship bet still stands, but the lab didn't stay quiet.

The next real test is the one the post already named: Google, projected ~August 12. The arcade will keep score either way.

Update (Aug 10, 2026): the silence rule scores a third hit

Ten more days, 13 more releases, and the figures above have been regenerated from the full dataset (240 releases through August 10). The board:

  • Alibaba's stall resolved exactly on pattern. The Aug 1 update flagged eleven days of silence against a ~1-day cadence and called it "winding up, not down." On August 1, Qwen3.8-Max landed — a flagship that debuted #1 on the text arena, alongside a 27B sibling. That's the third consecutive confirmation of this post's one durable rule: silence from a lab with an established cadence is itself a signal, and the longer it runs, the bigger the release. BFL, Anthropic, now Alibaba.
  • The open rebound found its strangest ally: Meta. Q3's open-weight share now sits at 52%, five of nine August releases are open — and on August 10 Meta shipped Muse Glimmer under Apache 2.0 while Zuckerberg and Alexandr Wang committed, by name, to open-weighting Muse Spark 1.2. Five days earlier this dataset had logged Meta's coding line as fully closed. If the ~50%-by-2027 call lands, the swing vote will have been the lab this post least expected.
  • The gating call got a bigger confirmation than "cyber tier" — gating moved upstream. The prediction was that trusted-partner access becomes a standard release tier. What actually happened: OpenAI announced it cannot rule out Critical cyber capability in Astra — an unreleased flagship — and is now gating training itself: isolated environments, weight encryption, paused internal workloads. The access-control tier didn't just become standard for releases; it reached into the lab.
  • "Orchestrators over monoliths" cashed out as the harness wave. The week of August 5–10, the capability story wasn't a model at all: three harnesses saturated ARC-AGI-3 (VISTA at 100%, Schema ~99%, Prime Agent 95.5%) without touching a single weight, and this blog had to open a whole new category to track them. "The unit of capability is shifting from a model to a system that routes models" was the June framing; by August the system doesn't even need to route multiple models to beat them.
  • Counterprogramming update. August 10 alone: Meta's Muse Glimmer and webAI's TwiL-LM family — both open-ish small models built to run on hardware you own — landed the same day, alongside a 6,500-word Zuckerberg essay. Multi-release days are now the norm, not the anomaly: Aug 4, Aug 5, and Aug 10 each carried two or three releases.
  • Still pending, now imminent: Google, projected August 12. Two days out as of this update. Anthropic's cadence math, refreshed, also points to ~August 14. If both ship big inside the week, the calendar itself will have become the most predictable thing about this industry.

The scorecard, quantified

Eighteen days of predictions is enough to grade with arithmetic instead of adjectives. The board, every call from the original post and its updates:

PredictionAs writtenOutcomeVerdict
BFL overdue → "FLUX.3 is big, or something corporate"flagged Jul 21FLUX 3, Jul 23 — multimodal pivotHit (2 days after flag)
Anthropic overdue, "the one to watch"projected Jul 12Opus 5, Jul 24 — debuted #1Hit (+12 days late)
Alibaba's "two-day standing appointment"next release ~Jul 23silence until Aug 1 (11-day gap)Miss as stated
Silence rule: quiet lab past cadence = big release brewingmeta-ruleBFL, Anthropic, Alibaba all resolved big3 for 3
Q3 lands north of 90fit said 3348 by day 41; projections 94–108Tracking
Open weights rebound toward ~50% by 2027quarterly trend32% → 36% → 52% (Q1→Q3)Early — at target 17 months ahead of the deadline
Cyber permanently gated3-for-3 then4 for 4, plus Astra's training now gatedHit, upgraded
DeepSeek flagship in Novembercadence projectionV4-Flash (a variant) Jul 31Push — shipped early, wrong class
Orchestrators over monolithsthesisthe harness waveHit — see the lever math below
World models break out 2027thesisno new entries since Cosmos 3 EdgeOpen
Google mid-Aug, Z.ai mid-Aug, xAI late SepprojectionsOpen (Google due in 2 days)

Four of the graded calls worth doing the actual arithmetic on:

The biggest miss is my own curve, and its failure is the most informative number on the page. The exponential fit's log-residuals over the last three quarters run +0.02, +0.38, +1.16 — 2026Q1 landed 2% above the fit, Q2 47% above, and Q3's pace is 3.2× above. Each quarter's overshoot roughly triples. That is not noise around an exponential; that's a process diverging from one, upward — super-exponential, or a regime change the functional form can't see. There's a pleasing symmetry here: the Skaling paper this blog covered the same day shows additive scaling laws failing exactly at the boundaries of their fitted grid. So does mine. A curve fit is a hypothesis, and the residuals are the referee. The Q3 punchline: the quarter consumed its entire fitted allocation of 33 releases by July 24 — day 24 of 92.

The Alibaba miss is a textbook base-rate error, and worth dissecting because I'll make it again otherwise. The "two-day standing appointment" came from the median of Alibaba's last five gaps — a recency window that had landed inside a release burst. Against the lab's full history, the picture inverts: across all 17 recorded inter-release gaps, the lifetime median is 35 days, and 13 of 17 gaps — 76% — are 11 days or longer. The "shocking" 11-day stall that the Aug 1 update treated as a signal was, by the lab's own lifetime distribution, a perfectly ordinary interval. The recency-window median wasn't measuring Alibaba's cadence; it was measuring one burst. Lesson, stated so it sticks: a five-sample median from an autocorrelated process is a description of the last burst, not a forecast — and the silence rule bailed out a bad point estimate, which is exactly the kind of rescue that shouldn't be counted as a win for the point estimate.

The orchestrator call can now be stated as a ratio. On ARC-AGI-3, holding the harness fixed and swapping the model a full generation (GPT-5.6 Sol → Claude Opus 5, both under ARC's official no-harness protocol) moves the verified score 7.8% → 30.2%: a 3.9× lever. Holding the model fixed and swapping the harness moves Opus 5 from 30.2% to 95.5% (Prime Agent) — 3.2× — and Sol from 13.3% to 78.3% — 5.9×. The scaffolding lever is now the same size as the model-generation lever, and on some models bigger. "The unit of capability is shifting from a model to a system" was a qualitative bet in July; by August it's a measured coefficient.

The cyber-gating call deserves its uncertainty stated. Four gated cyber models out of four sounds decisive, but under the 2026 continental base rates (73% of North American releases closed, 44% of Asian ones), the probability of all four landing closed by chance is 0.73³ × 0.44 ≈ 0.17 — suggestive, not conclusive, on the count alone. What upgrades it from statistical lean to structural fact is categorical, not numerical: Shieldstral shipping Apache 2.0 the same month (defensive capability open, offensive gated — the split is deliberate), and OpenAI moving the gate upstream of release entirely for Astra. The count is weak evidence; the asymmetry is strong evidence.

The running total: of eleven graded or gradable calls, six hit, one missed, one pushed, three remain open — and the miss taught more than the hits.

The bear case: reading the same curve as a bubble

Everything above treats the exponential as a growth story. It's worth taking seriously that the identical data supports a darker reading — because a process that grows faster than exponential, as this one now does, cannot be a steady state. Mathematically, a growth rate that itself grows implies a finite-time singularity: extrapolated naively, the release count goes vertical in finite time, which is impossible, so something else happens first. In the financial-bubble literature this is precisely the diagnostic — super-exponential growth with accelerating residuals is the signature of an unsustainable regime, not a healthy one. My +0.02 / +0.38 / +1.16 residual sequence is either a regime change or a blow-off top, and count data alone cannot tell you which.

So look at composition instead of count. Recomputing the dataset by lab structure per year:

YearReleasesDistinct labsEffective labs (1/HHI)New entrantsVariant-marker share
2022963.9622%
2023271411.61019%
2024412013.9824%
2025481610.6248%
20261153519.41649%

Two things jump out, and they cut against the comfortable reading. First, this is not consolidation — it's the opposite. 2026 is the most fragmented year in the dataset: an effective 19.4 labs sharing the releases, and sixteen first-time entrants after 2025 had nearly closed to entry with two. Second, half of everything shipped now carries a variant marker (Flash, mini, preview, Lite…) — double the 2022–2024 norm. The exponential isn't being driven by more frontier models; it's being driven by more labs shipping more SKUs of smaller things. An entry flood plus unit-size deflation is what the top of every entry boom has looked like — railway manias and dot-com included — because entrant count historically peaks precisely when capital is still cheap but differentiation is already gone.

The demand side doesn't double every four quarters

Here's the arithmetic problem underneath. Release count doubles roughly every 4.4 quarters. For the median release to hold its usage constant, total token demand would have to double on the same clock — and usage doesn't distribute evenly, it concentrates. Under a Zipf-like distribution (which inference traffic empirically resembles: ~61% of OpenRouter's open-model tokens go to a handful of Chinese models), the median model's share falls faster than 1/N as N grows. Doubling the number of models more than halves the median model's traffic. The release exponential is, mechanically, an attention-starvation machine for everyone below the top ten.

And for open models specifically, this blog has already documented three mechanisms that suppress usage below release share:

The cost side is already behaving like a commodity market

The revenue math meets a cost structure moving the other way. Frontier training runs cost more every generation, and the industry's pricing is doing what commodity pricing does: Luna cut to $0.20/$1.20, MiniMax undercut the video market ~3×, and Meta now sells Muse Spark 1.2 inference at a 90%+ discount in exchange for training rights — pricing below cost to acquire data, the textbook move of a market where the product itself no longer commands margin. Meanwhile the capability lever has visibly shifted to the harness layer, where the marginal cost of a breakthrough is an engineer-month rather than a cluster-year. When scaffolding worth 3–6× sits on top of models competing on price, the models are the commodity and the harness is the product. That is not a prediction; it's a description of August.

Four endings, and what each looks like in this dataset

The honest position is a scenario table with falsifiable signatures, not a forecast:

  1. Soft plateau. Demand growth and release growth converge; the log-residuals mean-revert toward zero within 2–3 quarters and quarterly counts stabilize somewhere in the 40–60 band. Signature: residual sequence turns negative; effective-lab count holds ~20.
  2. Shakeout. The 2026 entrant flood meets the margin math above and reverses, the way 2025's near-zero entry already previewed. Signature: new-entrant count collapses back toward 2, effective labs falls while variant share keeps rising — survivors shipping more SKUs each — and the HHI climbs for the first time in the dataset.
  3. Commoditization without collapse. Releases stop being events at all: the median gap goes below one day, labs move to continuously versioned channels, and this dataset's unit of account breaks. Signature: variant share through 60%, and the release count becomes formally unmeasurable — which the original post already flagged as the saturation endpoint.
  4. Continued super-exponential. Not an option for long, per the singularity argument — sustained, it forces one of the other three.

The uncomfortable summary: the bull reading and the bear reading agree on every number and disagree only on the noun. 115 releases from an effective 19 labs at 49% variant share is either "the most productive year in AI history" or "peak entry, pre-shakeout" — and the tiebreaker won't be in this dataset. It will be in the two numbers this dataset doesn't contain: tokens served per model, and dollars of margin per token. Release counts measure supply. Every collapse in history was a demand-side discovery.


Analysis: exclude the 7 non-release events (data through August 10, 2026; figures regenerated at each dated update), fit log-linear least squares on complete quarters only, medians for gap statistics; specialist counts use a keyword heuristic over model names and modalities; effective lab count is the inverse Herfindahl index of release share by lab, and variant share counts names carrying flash/mini/lite/preview/nano/turbo/edge-style markers or point versions.