François Chollet — creator of Keras, the ARC-AGI benchmark this blog has covered in depth, and now co-founder of Ndea — posted a conditional on the week's frontier-lab pacing proposals: he hopes they're genuine, but he's watching for two specific things that would tell him otherwise.
His framing is worth quoting at length, because the structure of the argument matters as much as the conclusion. He hopes the proposals "stem from a genuine concern for safety and a recognition of the potential risks posed by future models, rather than a strategic effort to consolidate power and permanently solidify the market dominance of a few top players." If genuine, he argues, oversight "must take a more democratic and accountable form, with both national and international components, similar to the Nuclear Regulatory Commission in the US and the International Atomic Energy Agency internationally." And then the test that gives the rest of it teeth: "If there is only one organization responsible for safety monitoring, and it happens to be staffed by the same people as the frontier labs, and it is perfectly aligned with them both by incentives and by ideology, then it would be indistinguishable from letting frontier labs self-regulate and self-certify their own safety standards."
From there, two concrete warning signs: "Calls to ban open-source AI. Open-source development currently serves as the only real counterweight to the dominance of frontier labs," and "Attempts to hinder non-frontier research. Models that are one or two generations behind the frontier — whose level of capability has already been deployed at scale — are empirically known not to pose safety risks." His conclusion: absent both signs, and with real international coordination underway, he'll believe the proposals are benevolent.
This isn't a reaction to any single document. It reads most directly against Amodei's "We Must Pace the Frontier" essay from the day before, but it's really a framework for judging the entire week — Amodei's essay, Bengio's causal account of why agents misbehave, Virkkunen's call for global AI rules, Wildberger's IAEA proposal cited in that same post. What makes it useful rather than just another opinion is that it's falsifiable, and this blog's own reporting already has data on both tests.
Test one: has anyone called to ban open-source AI?
Not this week, and not by Amodei. But there's a closer historical instance than a first read suggests, and it's worth being precise about how it lines up.
In July, this blog covered a real push to ban Chinese open-weight models outright, reported by Axios and escalated by Treasury Secretary Scott Bessent and OSTP director Michael Kratsios on distillation-theft grounds. It drew a fifty-signatory industry coalition — NVIDIA, Microsoft, Meta, Hugging Face, Mozilla, and dozens more — opposing it within days, and it collapsed to a narrower menu of procurement restrictions and Entity List threats rather than an actual ban. Dario Amodei himself, that same week, went on record: "Anthropic has never advocated for a ban on open-weights models," called open models without dangerous capabilities "a public good," and stated a blanket ban "is neither the correct remedy nor something we have called for."
That is Chollet's warning sign, narrower than his own framing: a ban push targeted at models from a specific country on national-security grounds, not a general call to ban open-source AI as a category, and it already failed to hold. Nothing in Amodei's new September essay revives ban language. What it does propose — a crackdown on "unauthorized distillation by companies in authoritarian countries" and continued chip export controls — is the same geopolitical argument from July, aimed at a competitor nation rather than at open licensing itself. Whether that distinction matters depends on where the line falls between "restricting China's access to compute" and "restricting open-weight release," and July's fight showed those two things get conflated easily under pressure, even if they weren't this week.
Test two: has anyone tried to hinder non-frontier research?
This is the harder one to score, because it turns on a question this blog flagged as unresolved back in July and that Amodei's new essay still leaves unresolved.
The July post's own conclusion, after the ban fight collapsed: "everyone now agrees on no ban, and the argument moves to compliance burden." Amodei's stated position then was "requiring safety testing of all sufficiently capable models, open and closed" — and the phrase doing all the work is "sufficiently capable." Nobody defined it. A testing mandate that applies only to genuine frontier releases is uncontroversial and consistent with Chollet's framework; the same mandate applied to a model one or two generations behind — Chollet's specific example of something "empirically known not to pose safety risks" — would be exactly the kind of non-frontier hindrance his second test flags.
Amodei's September essay reintroduces this same undefined threshold in a new form rather than resolving it: capability-triggered checkpoints, where a model that can, say, "defeat most common sandboxing methods" needs certified alignment evidence before release. That's framed entirely around frontier capability thresholds — nothing in the essay proposes testing requirements for older or smaller models, and it never mentions open-weight releases at all, favorably or otherwise. On the text, test two isn't tripped. But the essay also doesn't answer the question July left open, and the answer only becomes checkable once "sufficiently capable" gets written down as an actual number or capability benchmark rather than a phrase.
The self-regulation test lands on a name Amodei actually gave
Chollet's sharper and more specific claim is about organizational structure, not policy content: oversight that's "staffed by the same people as the frontier labs" and "perfectly aligned with them both by incentives and by ideology" is self-regulation with extra paperwork. Amodei's essay names a specific answer to who should do the overseeing — "embedded third-party evaluators (such as METR)" — which makes Chollet's test directly applicable rather than hypothetical.
METR's own history is on the record and worth stating plainly rather than insinuating. Its founder and CEO, Beth Barnes, is a former OpenAI alignment researcher who left in 2022 specifically to build an evaluation organization outside the labs, initially as ARC Evals under Paul Christiano's Alignment Research Center before spinning off as an independent nonprofit in December 2023. Its staff includes Ajeya Cotra, whose account of the OpenAI/Hugging Face incident this blog has quoted directly, a former Open Philanthropy safety grantmaker. And METR is not a hypothetical single point of oversight — it's already functioning as one in practice: it investigated OpenAI's Hugging Face breach and Anthropic's own cybersecurity eval incidents this summer, the same organization reviewing both of two competing frontier labs.
None of that is evidence of capture on its own, and it would be unfair to read it that way. METR was built by someone who left a lab specifically to create outside scrutiny, not by the labs themselves, and its reports on both OpenAI and Anthropic have been genuinely critical rather than exculpatory — this blog's own coverage of both incidents relied on METR findings that embarrassed the companies under review. But Chollet's test isn't about whether an evaluator is honest. It's about whether the same small, ideologically aligned community — people who move between frontier-lab safety teams, effective-altruism-adjacent funding bodies, and the evaluation organizations reviewing those same labs — ends up being the only structure checking all of them, regardless of any individual's good faith. METR reviewing both OpenAI and Anthropic this year is exactly the pattern his test asks a reader to notice, whatever conclusion you draw from noticing it. Amodei's proposal, notably, doesn't name a second or third organization, and doesn't address what happens if every frontier lab ends up embedding the same one.
The positive criterion is already partly in motion, on different terms than Amodei's
Chollet's last condition — real international coordination, not just national industry pacts — has a live candidate this blog already covered: German digital minister Karsten Wildberger's proposal for an IAEA-style AI safety body with mandatory incident reporting, raised the day before Amodei's essay, and explicitly citing the NRC/IAEA model Chollet names.
Where it differs from Amodei's own plan is instructive. Amodei's three-step structure puts national coordination among democratic-country labs second and genuine global coordination with China third and most conditional, explicitly bounded throughout by "keeping democracies' AI lead over autocracies as large as possible." Wildberger's version, and the EU's broader position this blog has tracked, treats international reporting as the more immediate ask, not the hardest-won final step gated behind preserving one bloc's competitive advantage. Chollet doesn't specify which sequencing he'd accept, but "real international coordination efforts" as his stated bar is closer to Wildberger's framing — an institution built to receive reports from anyone, not a coalition of allies agreeing to coordinate against a rival — than to the essay it was apparently written in response to.
Scored against Chollet's own three conditions as of today: no ban has been called for, and the one adjacent episode collapsed months ago with Amodei on record against it. No non-frontier hindrance has been proposed in writing, though the threshold that would create one remains undefined exactly where it was in July. And the single named evaluator sits inside a small, interconnected safety community reviewing every major lab at once — not proof of capture, but not the structural independence his own framework asks for either. Two of three conditions currently favor reading the proposals as genuine. The third is the one worth watching, because it's the one nobody involved has an incentive to fix on their own.