Moonshot AI announced Kimi K3 on July 16, 2026 — a 2.8-trillion-parameter Mixture-of-Experts model that the company says is now the largest open-weight model in the world. Early reporting pegs it as routing to 16 of 896 experts per token (around 50B active), though Moonshot hasn't officially confirmed the exact active-parameter count.
This post was written on launch day and updated three times as the story moved — an accusation, a weights release, a technical report. It's now organized around that chronology rather than around the model, because in K3's case the dates turned out to be the argument.
Where K3 comes from
K3 didn't appear from nowhere, and the release cadence behind it is worth seeing laid out:
| Date | Release | Step |
|---|---|---|
| Oct 2023 | Kimi Chat | 128K context — a world record at the time |
| Mar 2024 | 2M-character context beta | Long context as the house specialty |
| Jul 2025 | Kimi K2 | 1T open MoE — the pivot to open weights |
| Nov 2025 | K2 Thinking | Reasoning |
| Jan 2026 | K2.5 | Agent swarm, MoonViT-3D vision |
| Apr 2026 | K2.6 | 300-agent swarm, ties GPT-5.5 on SWE-Bench Pro at 80% lower cost |
| Jun 2026 | K2.7-Code | Coding specialist |
| Jul 2026 | Kimi K3 | 2.8T, 1M context, native multimodal |
That's five significant releases in the twelve months before K3, and a major architecture jump roughly every quarter. The K2 → K3 gap is almost exactly one year, during which the model went from 1.04T to 2.78T parameters and from 128K to 1M training context.
Two things follow. First, the compounding is visible in the artifacts, not just the version numbers: K2.5's MoonViT-3D became K3's MoonViT-V2, K2.6's agent swarms became K3's agentic RL, K2.7's coding work became the Kimi Code harness that K3's benchmarks run under. Second — and this becomes important later — a lab shipping at this rate leaves a public trail. That trail is what the July argument ends up turning on.
Two new pieces of architecture
K3 leans on two internally-developed components. Kimi Delta Attention (KDA) is a linear attention mechanism refining Gated DeltaNet with finer-grained gating over the model's recurrent memory, reportedly delivering up to 6.3x faster decoding at million-token contexts. Attention Residuals is described as a drop-in replacement for standard residual connections that gives consistent gains as the model scales. Combined with a 1M-token context window and native visual understanding, this is Moonshot betting on efficiency-at-scale rather than just adding parameters.
An always-on reasoner
Unlike models that toggle between a fast mode and a "thinking" mode, K3 ships with reasoning on by default — Moonshot calls it "thinking mode," and it's not something you switch on for hard problems, it's just how the model runs.
Where it actually lands
On GDPval-AA v2 — a benchmark spanning real-world tasks across 44 occupations and 9 industries — K3 scored 1,687, good for third place overall. Only Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8) finished ahead of it, with Claude Opus 4.8 (1,600) behind. In task-automation benchmarks it actually took first place in four of eight, including Automation Bench, SpreadsheetBench 2, and BrowseComp — areas where raw benchmark scores don't always predict who wins on messy, real work.
Pricing tells its own story: $15 per million output tokens, well above GLM-5.2's $4.40 or DeepSeek V4's $0.87, but still a fraction of what the closed frontier charges for comparable capability — roughly Opus-4.8-class performance at something closer to Sonnet-5-era pricing.
Not out yet, but close
(Written July 16.) The announcement landed just ahead of the 2026 World Artificial Intelligence Conference in Shanghai. Full model weights are scheduled to follow on July 27 — so for now, K3 is a benchmark story more than a downloadable one, but if the numbers hold once weights are public, it's a serious marker for how close open-weight models have gotten to the closed frontier.
They held, and the weights arrived early. What happened in between is the rest of this post.
Twelve days in July
The rest of this post covers the twelve days from K3's announcement to its weights going public — plus the one date before it that everything else hangs on. Laying it out by date matters, because the compression is the story: an entire policy cycle ran its course in under two weeks, and the model at the centre of it shipped anyway.
| Date | Event |
|---|---|
| Jul 1 | Anthropic's Claude Fable 5 becomes publicly available |
| Jul 16 | Kimi K3 announced; weights promised for Jul 27 |
| Jul 20 | Axios reports the administration is reviving a push to ban Chinese open-weight models |
| Jul 21 | Treasury Secretary Bessent: "watermarks of our U.S. large language models" on Chinese models; sanctions on the table |
| Jul 22 | OSTP director Kratsios publicly accuses Moonshot of distilling Fable to build K3 |
| Jul 24 | Industry coalition publishes "Open Weights and American AI Leadership"; Jensen Huang shares it in his first-ever post on X |
| Jul 26 | Weights ship — a day early |
| Jul 27 | Training infrastructure and PerceptionBench open-sourced |
Fifteen days from Fable 5's public release to K3's launch. Six from launch to a White House accusation. Ten from that accusation to Moonshot answering it with source code rather than a statement — they have still never issued one.
Update (July 22): White House accuses Moonshot of distilling Claude Fable
Six days after launch, the story around K3 got a lot more pointed. White House AI & Crypto czar Michael Kratsios posted that the US has information Moonshot AI distilled Anthropic's Fable to build K3. He was not describing the ordinary, legitimate kind of distillation used industry-wide to make smaller efficient models, but what he called a "sophisticated internal platform to conduct large scale distillation against U.S. models," built specifically to let Moonshot "quickly switch between multiple methods of access to avoid detection." He also alleged Moonshot acquired GB300-equipped servers and accessed GB300s in Thailand, likely to train its models — a detail that matters because GB300s are subject to US export controls.
This isn't the first version of this accusation. Anthropic itself said in February 2026 that Moonshot, alongside DeepSeek and MiniMax, had run large-scale distillation attacks against Claude. What's new is that it's now a formal US government allegation, made by name, about the specific model this post is about. If the claim holds up, it recasts K3's headline numbers — largest open-weight model, competitive with Fable-class performance, a fraction of the price — in a very different light. Moonshot hasn't publicly responded as of this update.
Within days, this accusation escalated into a full policy fight over whether the US should restrict open-weight AI models at all — a much bigger story than K3 itself.
Update (July 27): the weights landed — and so did the factory
Moonshot shipped, a day early. The weights went up on July 26 rather than the promised 27th, under the Kimi K3 License — MIT-style permissive terms, with a separate agreement required only if you run a Model-as-a-Service business above $20M revenue over any twelve months. That's a narrower restriction than "modified MIT" suggested: aimed at cloud resellers, not researchers.
Two corrections to what's above, now that the model card is public. Active parameters are 104B, not ~50B — 16 of 896 routed experts plus 2 shared, across 93 layers. And the GDPval-AA v2 figure in Moonshot's own table is 1,686, against Fable 5's 1,747 and GPT-5.6 Sol's 1,736. The architecture is 69 KDA layers interleaved with 24 Gated MLA, and the released weights are natively MXFP4 — quantization-aware training from the SFT stage onward, not post-hoc compression. For a 2.8T model that's the difference between "open" and "actually runnable."
But the weights are the least interesting thing Moonshot released this week.
They open-sourced the factory, not just the product
Alongside K3 came the infrastructure that built it — four separate pieces:
MoonEP is an expert-parallelism communication library, and it exists to solve exactly the problem a 896-expert model creates. When you route each token to 16 of 896 experts, the routing is never even: some experts get slammed, some idle, and because a training step finishes when the slowest rank finishes, your iteration time is set by the hottest GPU. MoonEP's trick is dynamic redundant experts — it plans a small number of duplicate experts online from the current router outputs and prefetches them, so every rank receives exactly the same token count no matter how skewed the routing gets. Their benchmarks against DeepSeek's DeepEP v2 show the payoff: DeepEP's iteration time climbs steadily as imbalance grows and eventually OOMs outright, because shifting activation shapes fragment GPU memory. MoonEP stays flat, because the shapes are static by construction.
AgentENV is the one that explains K3's agentic scores. It's a platform for running Firecracker microVMs at scale, and the README states plainly that it powers "agentic RL training for Kimi K3." Environments boot or resume in under 50 ms and pause in under 100 ms; a running environment can fork into multiple independent sandboxes for parallel rollouts; images load on demand via overlaybd so the image set can exceed local disk. If you want to train a model by letting it actually use a terminal a few hundred million times, this — not the model code — is the hard part. It exposes an E2B-compatible API, so you can point the standard E2B SDK at it unchanged.
FlashKDA is the CUTLASS kernel implementation of Kimi Delta Attention, and its date is the most interesting thing about it. It's been open since April 22, 2026 — three months before K3 launched, and more than two months before Claude Fable 5 was publicly available. It's upstreamed into flash-linear-attention, so anyone can call chunk_kda and get Moonshot's kernels. Hold onto that timing; it comes back below.
PerceptionBench is a benchmark rather than infrastructure, and it's the most self-aware of the four. Rather than another holistic vision eval, it isolates atomic perception: the authors diagnosed the earliest failure point in frontier model responses across 42 existing benchmarks, built an error taxonomy from that, and wrote 3,000 verified questions each targeting exactly one of ten perceptual capabilities, with difficulty coming from seeing rather than reasoning.
The benchmark they published and came second on
| # | Model | Overall | Hallucination |
|---|---|---|---|
| 1 | GPT-5.6-Sol | 59.7 | 26.9 |
| 2 | Kimi K3 | 58.5 | 41.7 |
| 3 | Claude-Fable-5 | 57.2 | 45.0 |
| 4 | Gemini-3.1-Pro | 56.2 | 40.6 |
| 5 | GPT-5.5 | 55.8 | 34.7 |
The headline finding is a negative one: no model reaches 60%. Atomic visual perception — counting, depth, localization, telling whether two things are the same colour — is substantially unsolved across sixteen frontier models, and the paper notes that similar overall scores hide sharply divergent capability profiles. The hallucination column makes the point: GPT-5.6 Sol wins overall while scoring 26.9 on perception-related hallucination, where K3 scores 41.7 and Fable 5 leads at 45.0. Ranking these models by one number was always hiding something.
Publishing a benchmark your own flagship comes second on is a real credibility move, and worth saying so. It isn't a free pass — it's Moonshot's benchmark, built from Moonshot's taxonomy, graded by an LLM judge — but they plainly didn't construct it to win.
Does this strengthen the release, or not?
Mostly yes, and in a way the benchmark tables don't capture.
Weights are a snapshot; infrastructure is a capability. A 2.8T checkpoint is enormously useful and completely inert — you can run it, fine-tune it, and that's the end. MoonEP and AgentENV are reusable by anyone training a large sparse MoE or doing agentic RL, including Moonshot's competitors. That's a materially more open posture than shipping weights alone, and it's the opposite of what a lab does when its advantage is a secret.
It also lands, deliberately or not, in the middle of the distillation fight. You do not build a load-balancing EP library and a Firecracker RL harness if your method is copying someone else's outputs. Both artifacts are evidence of an organization doing expensive pretraining and RL infrastructure work at frontier scale, and both are now inspectable. That's a stronger rebuttal of the Kratsios accusation than any statement Moonshot could have issued — which they still haven't.
The strongest evidence is a timestamp
Set the artifacts aside for a moment and just look at the calendar.
Claude Fable 5 became publicly available on July 1. K3 launched on July 16. That's the two-week window skeptics have pointed at from the beginning — a very short time in which to covertly distill a frontier model, train a 2.8T MoE on its outputs, and ship it.
But the sequencing goes back much further than that fortnight. FlashKDA — the kernel implementation of the attention mechanism K3 is built on — was public on April 22, two and a half months before Fable 5 existed as a product anyone could query. The KDA and Attention Residuals research predates it further still. The 2.5× scaling-efficiency curves in the technical report describe a training run, and a 2.8T model's pretraining is measured in months, not the fifteen days available.
So the timeline constrains what the accusation can coherently mean. K3's architecture demonstrably predates the model it's alleged to have been copied from. Any surviving version of the claim has to be about post-training data — outputs harvested to shape behaviour — layered onto a model whose foundations were already laid and publicly documented. That's a much narrower charge than "distilled Fable to build K3," and it's one the release genuinely doesn't answer.
Which is the honest limit here. Training infrastructure demonstrates capability, not data provenance. MoonEP proves Moonshot can train a huge MoE efficiently; it says nothing about what was in the corpus. Timestamps prove the architecture wasn't copied; they don't prove the post-training data was clean. The releases make the accusation as stated much less plausible without disproving a narrower version of it.
Two other things worth noticing. MoonEP's benchmarks run on H20 — the export-control-compliant chip — as do several of K3's own evaluations, which is a quiet counterpoint to the claim that K3 depended on smuggled GB300s. And AgentENV ships with a blunt warning that it has no authorization support and must not be exposed to a public network, which is a useful reminder that "open-sourced alongside a frontier model" doesn't mean "production-hardened."
The honest summary: the weights made K3 downloadable, but the infrastructure is what makes it reproducible-in-principle. For a release whose central contested question is "did they really build this," shipping the tools that built it is the most persuasive answer available.
Update (July 27): what the technical report adds
Published with the weights, the full 47-page report fills in what the model card only gestured at — and, read against the calendar above, most of what it describes could not have happened in July.
The scaling story is efficiency, not size
K3 is 2.7× K2's parameter count, but Moonshot's claim is that the architecture changes bought an approximately 2.5× gain in scaling efficiency — measured as fitted scaling-law curves on held-out out-of-distribution validation data, not a benchmark delta. The full comparison:
| Kimi K2 | Kimi K3 | Δ | |
|---|---|---|---|
| Layers | 61 | 93 | +52% |
| Total parameters | 1.04T | 2.78T | +167% |
| Activated parameters | 32.6B | 104.2B | +220% |
| Routed experts | 384 | 896 | +133% |
| Experts active / token | 8 | 16 | +100% |
| Shared experts | 1 | 2 | +100% |
| Attention heads | 64 | 96 | +50% |
| Training context | 128K | 1M | 8× |
| Attention | 61 MLA | 69 KDA + 24 MLA | hybrid |
| Activation | SwiGLU | SiTU-GLU | — |
| Hidden dimension | 7,168 | 7,168 | unchanged |
The unchanged row is the interesting one. K3 grew deeper and sparser, not wider — more layers, more experts, more active experts, same hidden dimension. They also re-ran scaling-law searches for batch size, learning rate and tokens-per-parameter rather than inheriting K2's, and found cosine decay beat Warmup-Stable-Decay — with the caveat, stated carefully, that the two schedules have such different optimal hyperparameters that comparing them under a shared setting unfairly favours whichever one the shared setting happens to suit.
One more architectural note worth flagging: vision is trained jointly from the start, with visual and text tokens interleaved under a single next-token objective — not a vision encoder grafted onto a finished language model via post-hoc alignment. That's the harder path, and it's consistent with K3's unusually strong showing on OmniDocBench and Video-MME.
K3 designed a chip, and the RTL is public
The case-studies section is the part that will get argued about. Two artifacts, both produced by K3 rather than used to build it, and both open-sourced:
MiniTriton — a compact Triton-like GPU compiler K3 wrote: custom tile-level Python frontend, warp-level MLIR annotation layer, PTX code generation, plus a dual-mode tensor library with reverse-mode autograd and NCCL distributed primitives. On an L20 it beats PyTorch eager and torch.compile in geometric mean over its benchmark suite, and its from-scratch tensor-core matmul reaches roughly 90% of measured machine roof. It trains a GPT model end to end with gradients matching torch autograd to within torch's own fp32 rounding error.
nano-kpu — an inference-chip prototype. In a single 48-hour autonomous run with Kimi Code, K3 built, optimized and verified a chip using open-source EDA tools against the Nangate45 standard-cell library. Inside a 4 mm² area budget it closes timing at 100 MHz and hits an RTL-simulated decode throughput of over 8,700 tokens/s, with 1.46M standard cells, 0.277 MiB of SRAM and an INT4 MAC array with fused dequantization.
Both are proofs of capability rather than products, and "RTL-simulated" is doing real work in that sentence — no silicon exists. But an autonomous 48-hour run producing a timing-closed design is a substantially more concrete agentic claim than a leaderboard position, and unlike a benchmark score, the output is inspectable.
The report uses the word "distillation" — carefully
Given the accusation above, this is worth being precise about. K3's post-training pipeline is explicitly built on distillation: SFT for a cold start, then RL to develop domain-specialized experts at different reasoning-effort levels, then Multi-Teacher On-Policy Distillation (MOPD) to consolidate those experts back into one model.
Every teacher there is Moonshot's own. This is self-distillation across their own specialist checkpoints — the ordinary, industry-standard kind, and the same technique Kratsios explicitly carved out as legitimate. It is not evidence for or against the allegation, which concerned distilling Anthropic's model, and the report is silent on that question. It contains no decontamination or data-provenance section at all.
What the report does supply is a detailed account of an organization doing very expensive original work: per-head Muon orthogonalization, KDA Context Parallelism derived from first principles with proofs in the appendix, a custom EP library with its own upper-bound proof. None of that is what copying looks like. None of it settles what was in the corpus either.
They say plainly that they lose
The abstract's own framing: K3's overall performance "still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol," while consistently beating everything else in their suite. For a launch document, naming the two models you lose to in the abstract is unusually direct — and it matches the benchmark tables, where K3 leads on agentic and retrieval work and trails on HLE and GDPval.
What the clock says
Three different clocks run through this story, and they disagree in a useful way.
The release clock is fast and getting faster. Five significant models in the twelve months before K3, a major architecture jump each quarter, 1.04T to 2.78T parameters in a year. On the day K3's weights landed, its immediate predecessor was six weeks old.
The policy clock is faster still, and less productive. A ban proposal, a Treasury broadside, a White House accusation and a 200-company counter-coalition, start to finish in eight days — with, at the end of it, no bill, no executive order, and the weights on Hugging Face anyway. The one thing everyone agreed on is that the files, once downloaded, cannot be recalled. The policy cycle finished faster than the release cycle and changed nothing about it.
The research clock is the slow one, and it's the one that actually decides things. KDA, Attention Residuals, the scaling-law work, FlashKDA in April — these took quarters, and they're why K3 exists. They're also, incidentally, why the fastest accusation of the summer doesn't fit the evidence: you cannot compress a year of architecture work into a two-week window, and Moonshot's habit of shipping in public left a timestamped trail proving it didn't have to.
That's the durable lesson from a month that mostly generated noise. In a field this fast, shipping continuously in the open is not only a distribution strategy — it's an alibi. A lab that publishes its kernels in April doesn't have to argue about what it built in July. The record argues for it.
Whether the next step is K3.5 in October is, on this cadence, less a question than a schedule.