2026-08-18

Claude Designed a Protein Binder That Beat 245 Human Entries — and Two Outside Labs Verified It

AIScience🌍 North America

Anthropic published two results from what it's calling Claude Science: an autonomous protein-binder design campaign, and a chemistry-lab analysis task. The protein design result is the one worth reading closely, because it's independently verified rather than self-graded — a distinction this blog has flagged as the difference between a vendor claim and evidence, most recently when OpenAI's own math claims needed a mathematician's vouching to be taken seriously after an earlier overclaim, and here that role is filled by two outside contract labs running the actual wet-lab tests. One caveat sits above everything below it: this is a self-published research post and a pair of accompanying technical reports, not a peer-reviewed paper. Independent wet-lab testing and a full data release are real, checkable steps toward scientific credibility, but they are not a substitute for outside experts formally vetting the methodology and conclusions — that step, if it happens, hasn't happened yet.

What got tested, and by whom

Claude — specifically Mythos Preview and Opus 4.8 — was set loose on de novo protein binder design: creating small proteins from scratch that latch onto a specific target, the step that historically costs a protein engineer weeks to months per target. Across 15 targets, Claude produced 1,320 designs, and Adaptyv Bio and Twist Bioscience — independent external contractors, not Anthropic — physically synthesized and tested them in the lab. That's the load-bearing detail: hit rate and affinity numbers here come from someone with no stake in the result confirming or refuting them, not from Anthropic's own benchmark.

The headline hit rates: 22.6% for Opus 4.8 and 26.7% for Mythos Preview when both designed against all 15 targets simultaneously in a single 48-hour session, rising to 35.1% for Mythos Preview when it worked one target at a time across parallel 24-hour sessions. Anthropic's stated baseline for the field is 10-15%. Worth flagging the single-target comparison isn't quite apples to apples, though: single-target mode gave Mythos Preview 24 dedicated hours per target, versus 48 hours split across all 15 in multi-target mode — meaningfully more session time and compute per target, not purely a change in strategy. The improvement is real, but some of it is resourcing, not just focus.

Claude didn't fold proteins itself — it orchestrated the tools that do

A different, more favorable take came from Y Combinator partner Ankit Gupta on X: "neat thing about this is all the protein design tools claude used are open source... as OSS protein design tools get really good, we're going to have the intelligence to design binders increasingly democratized bc tools like claude are smart enough to use them." That's directly confirmed by Anthropic's own description of the campaign: "Claude did this by operating publicly available specialist protein design and co-folding models that the field already uses." Claude's own design pipeline, per the campaign data, ran through seven structure-design tools — PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, RFdiffusion, and Proteina-Complexa were the named ones with meaningful volume — feeding into sequence-design tools (SolubleMPNN handled the large majority of designs, with SolubleCaliby, native co-design, and ProteinMPNN used less often), then cycled through rounds of in silico optimization before anything went to the wet lab. FreeBindCraft-originated designs had the highest hit rate of any structure-design method with substantial volume, a specific and checkable detail rather than a vague "the pipeline worked well."

One detail in Gupta's post we couldn't verify against the data available to us: he names ESMFold2 specifically as a tool this campaign used. Anthropic's materials describe ten structure predictors used for co-folded predictions without naming all of them individually in what we could review, so that specific claim is plausible but unconfirmed here. The broader point — that Claude's contribution is orchestration of existing open-source specialist models rather than an end-to-end model doing the folding and generation itself internally — is accurate and worth sitting with alongside the "grad student project" critique above: whether you read this campaign as evidence of Claude's own scientific capability or as evidence that a capable general model can now direct an existing open-source toolchain competently and at scale depends on which of those two things you think is the more interesting claim. Both readings are consistent with the same data.

A pushback worth weighing: how novel were the targets?

A verified X account under the name Melinda B. Chu, describing herself as a researcher, posted a pointed critique of the announcement: that this reads as "a grad student or Post-doc project... first-year grad student or college student project," that the targets were "primarily known targets AND Adaptyv Bio already knew how to synthesize them b/c they were known targets," and asked why the result is being treated as a major accomplishment rather than routine work. We couldn't independently verify her professional credentials or affiliation — X is unreachable from this environment, so this is relayed from a screenshot rather than the primary source — but the factual part of her claim is checkable against Anthropic's own text, and it's worth separating what's verifiable from what's a judgment call.

On target selection, Anthropic's own methodology confirms the core of her point: "We began our protein design campaign by selecting multiple targets that are commonly used in protein design benchmarks, including all of Adaptyv Bio's BenchBB. Because these targets have been studied extensively, we can compare our results against published hit rates and affinities." Of the roughly 15-16 targets used, Anthropic names only 15-PGDH and GDF-8 as targets chosen specifically for not having established prior art, "to ensure Claude was able to design against targets without drawing upon pre-recorded successes in its training data." So yes — most targets were established benchmarks, and Anthropic says so directly, framing it as a deliberate choice for comparability rather than something concealed. Her inference that Adaptyv Bio already had synthesis and assay protocols worked out for these targets is plausible — several of them are targets Adaptyv had already run public competitions on — but that specific claim isn't stated outright in Anthropic's materials, so it's a reasonable extrapolation rather than a confirmed fact.

In a follow-up post, Chu sharpened the point: "This is repeating an open contest from a few months ago. The person that won that contest didn't say they were curing all diseases." That specific claim — that RBX1 was already the subject of a public contest Anthropic's campaign then re-ran — is, again, confirmed directly by the post above: the RBX1 comparison is against that same Adaptyv Bio competition. Her broader point is a mismatch in scale between the claim and the evidence: Anthropic's own public messaging, including Dario Amodei's stated ambition to cure most human disease within roughly 5-10 years, is far grander than "matched a known benchmark's winning entry on one target out of fifteen." That gap between a modest, specific experimental result and the scale of the ambition it gets cited in support of is a fair thing to hold in mind — a narrow win on a repeated benchmark is real evidence of something, but it's not itself evidence of curing disease, and nothing in the announcement claims otherwise directly; the connection is made by the surrounding narrative, not by this specific result.

Where her framing is a judgment call rather than a checkable claim is the "grad student project" characterization itself — that's an assessment of significance and difficulty, not a fact this post can verify or refute. What the post's own numbers offer as context: on RBX1, one of the targets in question, Claude's design beat the winning entry among 245 submissions to a public competition — submissions that presumably included exactly the population of researchers, including grad students and postdocs, her comment invokes as the appropriate comparison. Using well-studied targets does make the comparison more legible (there's a known baseline to beat), and it also means this campaign is not evidence of Claude succeeding on genuinely novel biology — a narrower and more accurate claim than "Claude accelerates protein design" might suggest to a casual reader. Both things can be true at once, and the post above should be read with that in mind.

The RBX1 result is the specific one to remember

Against RBX1, a protein involved in targeted protein degradation, Mythos Preview in single-target mode hit 40% — and Adaptyv Bio had run a public competition on the same target where human entrants achieved a 3.7% hit rate across 245 submitted designs. Claude's top-ranked RBX1 design outperformed that competition's winning entry on affinity. That's a specific, falsifiable claim against a specific, dated public leaderboard — the kind of comparison this blog has repeatedly asked for and rarely gotten when a lab claims to beat "the field."

Where the two models disagreed, and Anthropic doesn't know why

The most candid line in the report concerns TNFα, a clinically important, structurally difficult target (it's the basis for Humira). Opus 4.8 designed working binders — including some that bind human, cynomolgus monkey, and mouse TNFα simultaneously, useful for animal studies — while Mythos Preview, the model that outperformed Opus 4.8 everywhere else in this campaign, failed on it entirely. Anthropic's own sentence: "We're not sure why Opus 4.8 was successful on this target and Mythos Preview was not." That's a real admission of a capability being non-monotonic across their own model line on a specific task, stated plainly rather than smoothed into a composite "our models are strong at protein design" claim. It's also a useful caution against reading any single hit-rate number as a stable ranking: on this one target, the "worse" model won outright.

Claude also produced 15 confirmed binders using β-sheet structures across six targets — a harder, misfolding-prone motif most computational design defaults to avoiding in favor of α-helix bundles — and the report is equally direct about failure: zero of 90 designs against maltose-binding protein were confirmed to bind (one showed a weak, reproducible signal), and its BBF-14 binders, while real, only reached modest sub-micromolar affinities against a target specifically chosen for being adversarially hard to design against.

The chemistry task: no autonomy claim, just a stopwatch

The second experiment used Opus 5, the generally available model — not a gated research variant — and tested something narrower: given a contract lab's raw NMR and LC-MS instrument files and a two-sentence prompt, could it reproduce the lab's own analysis? It returned results in 23 and 19 minutes, with hydrogen counts within 0.08 ppm of the lab's figures and a purity reading of 96.4% against the lab's 96.33%. Two details make this more than a speed claim: Claude independently proposed the same heavy-water follow-up experiment the lab had run three days after its own first measurement, and when it ran that follow-up data, it caught and corrected its own earlier overstatement — its first pass claimed four flagged peaks had all disappeared; its self-check showed only two had. That's a specific, visible instance of self-verification catching a real error, not just a faster final answer.

The data release backs the claim up further than the blog post does

Anthropic followed the announcement with a full data release — 1,440 designs in total (900 from Mythos Preview, 540 from Opus 4.8) across 16 targets, one more than the blog post's headline 15: the extra target, mature GDF-8, is included with its designs and provenance but its wet-lab measurements are explicitly excluded because the antigen aggregated and bound assay surfaces non-specifically, an inconclusive result reported as inconclusive rather than dropped without comment. The two-CRO verification is more rigorous than "two labs agreed": Adaptyv Bio and Twist Bioscience ran genuinely different assay methodologies on the same designs — cell-free expression with the design itself immobilized for SPR/BLI kinetics at Adaptyv, Fc-fusion expression with a six-point antigen titration under capture SPR at Twist — so a binder counted as confirmed had to clear two different experimental setups, not just two different buildings running the same protocol.

What's released goes well past the numbers in the announcement: raw sensorgrams and report images from both vendors per design, the final per-design comparison and verdict between the two CROs, co-folded structure predictions from ten different structure predictors at five seeds each (113,550 predicted structures with full PAE matrices), step-level design provenance, and — genuinely unusual — the exact campaign prompts, kickoff messages, and the full external resource corpus (cited papers, ProteinBase collections, method papers) Claude had access to during the campaign, each with its own redistribution/licence notes. Data and documentation are CC BY 4.0. That's a reproducibility package closer to what an independent academic group would need to actually re-run the analysis than what usually accompanies a lab's capability announcement, and it strengthens the credibility case in the same direction as the two-vendor validation itself: nothing here has to be taken on Anthropic's word alone.

The dual-use gate, stated plainly

Anthropic is explicit that protein design and "other dual-use research biology capabilities" remain withheld from Claude Fable 5's general access, available only through a "trusted access program" it says is coming soon — the same shape of gating this month's cyber-capability coverage has tracked repeatedly, applied here to biology instead of cyber, and for the same stated reason: autonomous research capability is dual-use, and the wet-lab campaign described above ran on Mythos Preview and Opus 4.8 specifically because the most capable models are still blocked from this class of task pending that access program.

What to expect next

  • Watch for the promised follow-up characterization. Anthropic states outright it intends further work to confirm these hit rates and affinities — a real commitment to re-checking its own headline numbers rather than letting a first result stand as final.
  • Watch whether the TNFα asymmetry recurs on other targets. One unexplained model-capability inversion is a data point; a second would suggest something structural about how these models approach binder design differently, not noise.
  • Watch the trusted-access program's actual terms, since that's what determines whether outside researchers — not just Anthropic's own team — can independently reproduce any of this on a model more capable than the ones used here.
  • Watch for peer review. Nothing here has gone through a journal or formal external review yet; the technical reports and open data make that step possible, but publication and independent critique remain the actual test of whether these results hold up under scrutiny.

References: Anthropic — How Claude is accelerating protein design and analytical chemistry · Hugging Face — Anthropic/claude-protein-binder-design data release · Melinda B. Chu on X — further discussion · Anthropic on X · related coverage: OpenAI's Math Claims Have a Credibility Problem · The Science Vertical: One Big Slice and Some Crumbs · OpenAI Paused Its Biggest Training Run Because One Model Might Be Too Good at Cyber · Claude Opus 5 · Frontier Arcade: trends & predictions