2026-08-03

Japan's Sovereign AI Strategy Is to Not Train a Model

AIOpen SourcePolicy🌍 Asia

Every sovereignty story this blog has covered so far involved somebody training a model. Korea's Dopamo programme spent a training run on A.X K2. Germany's Soofi S did too. France has Mistral. The assumption underneath all of them is that owning a national model means building one.

Sakana AI's Namazu is the argument that you don't have to. It isn't a foundation model at all — it's a post-training pipeline applied to other labs' open-weight models, to adapt them to a specific country. And its most-cited result is the one that makes the whole strategy legible: it takes a Chinese open model that refused 72% of questions about politics, history and diplomacy, and drops that refusal rate to roughly zero.

What Namazu actually is

Sakana — the Tokyo lab this blog has covered mostly for its Fugu orchestrator line — announced the Namazu series on March 24, 2026 as an alpha, alongside a free consumer chat product, Sakana Chat, with web search built in. The naming is a joke that lands better in Japanese: namazu is the giant catfish of Japanese folklore, the one that causes earthquakes when it thrashes.

The alpha shipped as three variants, and the list is the whole thesis:

  • Namazu-DeepSeek-V3.1-Terminus — a Chinese open model
  • Llama-3.1-Namazu-405B — an American open model
  • Namazu-gpt-oss-120B — OpenAI's open-weight release

Three different countries of origin, one post-training recipe, all adapted for Japanese language, cultural norms and factual coverage. Sakana isn't picking a base model; it's demonstrating that the base model is interchangeable.

The 72% number, and why it's the interesting one

The headline result is about censorship, and it's worth being precise about what was measured. On a benchmark of politically, historically and diplomatically sensitive questions relevant to Japan, DeepSeek-V3.1-Terminus refused to answer 72% of prompts. After Sakana's post-training, the Namazu variant's refusal rate fell to approximately 0% — while, per Sakana's evaluation, improving factual accuracy and neutrality on those same topics.

The other half of that claim is the one that would normally be the catch: does stripping out refusals wreck the model? Sakana reports near-parity with the base models on the standard capability suite — AIME'25, MMLU-Redux, GPQA Diamond, LiveCodeBench, IEval. Capability retained, refusals removed, factual coverage improved.

Note what this is not. It isn't a jailbreak, and it isn't an "uncensored model" in the edgelord sense. It is, technically, the same operation Thinking Machines' safety framework runs deliberately as a red-team test: adversarially fine-tuning refusals out of a model to see what dangerous capability that exposes. The Namazu case is narrower and more defensible: the refusals being removed are ones a Chinese model applies to questions about the Senkaku Islands or wartime history, which are not safety refusals in any meaningful sense — they're political alignment inherited from the training jurisdiction. That distinction matters, and the fact that both operations use the same technique is precisely why the safety debate around open weights is hard.

The strategy this makes possible, and what it costs

Put the sovereignty options side by side and Namazu's economics are startling.

StrategyExampleWhat it takes
Train a sovereign frontier modelA.X K2 (Korea, 688B)A full training run, state backing
Train a sovereign small modelSoofi S (Germany, 31.6B)A smaller run, still from scratch
Post-train someone else's open weightsNamazu (Japan)A post-training pipeline. No pretraining.

Sakana's own framing for this is sovereign AI — the idea that instead of using foreign AI as-is, each country readjusts it to its own characteristics, and secures AI sovereignty that way. It's the cheapest entry into the sovereignty game by an enormous margin, and it only exists because the weights are open. You cannot do this to GPT-5.6 Sol or Claude Fable 5 at any price.

That's the part worth sitting with. The six-rung openness ladder is usually read as a question about what you may build and sell. Namazu shows a different consequence of rung two: open weights are what let a country fix a model's politics. Japan didn't need Beijing's permission to delete DeepSeek's refusals, and it didn't need Washington's to adapt Llama. It needed a download.

The irony at the centre of it

Now put this against the week Washington flirted with banning Chinese open weights, on the argument that Chinese models carry Chinese values and shouldn't be running in allied infrastructure.

Namazu is the counterexample, executed by a US-allied democracy. Japan's answer to "this Chinese model is censored" was not to ban it — it was to download the weights and post-train the censorship out, in public, with a published benchmark. And the direction of travel since March makes the point harder: reporting indicates the Namazu API is built on Moonshot's Kimi K2.6, refined with Sakana's proprietary Japanese data, with Sakana Translate (a Japanese–English–Chinese translation, proofreading and Q&A tool) shipping on it in early July.

So a Japanese company is now running commercial products on Chinese open weights, having first demonstrated it can neutralize the alignment concern that made those weights politically controversial. Whatever the right policy is, "the values are baked in and can't be removed" is now an empirical claim with a published counterexample against it.

This also reframes the red-tsunami reading of Chinese open weights. Moonshot's models aren't just cheap capability — they're becoming substrate. When a foreign lab builds its sovereign national offering on your weights, you've achieved something an API business can't buy, and it costs Moonshot nothing per Japanese query.

The caveats, which are real

Every number here is Sakana's own. The 72%-to-zero refusal result, the neutrality improvement, and the capability-retention claim all come from the lab's own evaluation on its own benchmark, with no independent reproduction I could find — the same caveat this blog applies to Alibaba's benchmark tables and everyone else's. "Neutrality" on contested historical questions is also a construct that a Japanese lab defining a Japan-relevant benchmark is not a disinterested party in measuring. The result is important either way, but it's a vendor claim.

I also could not reach Sakana's own pages directly while writing this — sakana.ai returned 403 from this environment, as several lab domains have all week — so the specifics above come from contemporaneous coverage of the March announcement and subsequent reporting rather than the primary post. The alpha lived at /namazu-alpha/; the URL is now /namazu/, which suggests the programme has consolidated past its alpha framing, but I could not confirm what changed.

What to expect next

  • "Sovereign post-training" becomes a category with more than one entrant. The strategy is cheap, the base models are free, and every mid-sized country with a language and a political history has the same problem Japan does. Expect an Indian, Brazilian or Gulf equivalent within a year.
  • Refusal-rate-by-jurisdiction becomes a published metric. Sakana has demonstrated it's measurable and fixable. Once one lab publishes the number for a rival's model, everyone has to.
  • The open-weights ban argument gets harder, not easier. The strongest case against Chinese open models was value alignment. Namazu doesn't refute the concern, but it demonstrates a remedy that doesn't require export controls — and remedies tend to beat bans in policy fights.

References: Sakana AI — Namazu · Sakana AI — Namazu alpha announcement · ITmedia — Namazu and Sakana Chat launch · Impress Watch — adapting even DeepSeek to Japanese specifications · GIGAZINE — Sakana Chat launch · MarkTechPost — Sakana Translate · StartupHub — Namazu adapts global models for Japan · related coverage: A.X K2 · Soofi S · the open-weight ban fight · the openness ladder