Anthropic published two results from what it's calling Claude Science: an autonomous protein-binder design campaign, and a chemistry-lab analysis task. The protein design result is the one worth reading closely, because it's independently verified rather than self-graded. This blog has flagged that distinction as the difference between a vendor claim and evidence, most recently when OpenAI's own math claims needed a mathematician's vouching to be taken seriously after an earlier overclaim. Here the verifying role is filled by two outside contract labs running the actual wet-lab tests. One caveat sits above everything below it: this is a self-published research post and a pair of accompanying technical reports, not a peer-reviewed paper. Independent wet-lab testing and a full data release are real, checkable steps toward scientific credibility, but they are not a substitute for outside experts formally vetting the methodology and conclusions. That step, if it happens, hasn't happened yet.
What got tested, and by whom
Claude — specifically Mythos Preview and Opus 4.8 — was set loose on de novo protein binder design: creating small proteins from scratch that latch onto a specific target, the step that historically costs a protein engineer weeks to months per target. Across 15 targets, Claude produced 1,320 designs, and Adaptyv Bio and Twist Bioscience — independent external contractors, not Anthropic — physically synthesized and tested them in the lab. That's the load-bearing detail: hit rate and affinity numbers here come from someone with no stake in the result confirming or refuting them, not from Anthropic's own benchmark.
The headline hit rates: 22.6% for Opus 4.8 and 26.7% for Mythos Preview when both designed against all 15 targets simultaneously in a single 48-hour session, rising to 35.1% for Mythos Preview when it worked one target at a time across parallel 24-hour sessions. Anthropic's stated baseline for the field is 10-15%. Worth flagging the single-target comparison isn't quite apples to apples, though: single-target mode gave Mythos Preview 24 dedicated hours per target, versus 48 hours split across all 15 in multi-target mode — meaningfully more session time and compute per target, not purely a change in strategy. The improvement is real, but some of it is resourcing, not just focus.
Claude didn't fold proteins itself — it orchestrated the tools that do
A different, more favorable take came from Y Combinator partner Ankit Gupta on X: "neat thing about this is all the protein design tools claude used are open source... as OSS protein design tools get really good, we're going to have the intelligence to design binders increasingly democratized bc tools like claude are smart enough to use them." That's directly confirmed by Anthropic's own description of the campaign: "Claude did this by operating publicly available specialist protein design and co-folding models that the field already uses." Claude's own design pipeline, per the campaign data, ran through seven structure-design tools: PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, RFdiffusion, and Proteina-Complexa were the named ones with meaningful volume. Their output fed into sequence-design tools — SolubleMPNN handled the large majority of designs, with SolubleCaliby, native co-design, and ProteinMPNN used less often. Everything then cycled through rounds of in silico optimization before anything went to the wet lab. FreeBindCraft-originated designs had the highest hit rate of any structure-design method with substantial volume, a specific and checkable detail rather than a vague "the pipeline worked well."
One detail in Gupta's post goes beyond what Anthropic published: he names ESMFold2 specifically as a tool this campaign used. Anthropic's materials describe ten structure predictors used for co-folded predictions without naming all of them individually, so that specific claim is plausible but unconfirmed. The broader point is accurate: Claude's contribution is orchestration of existing open-source specialist models, not an end-to-end model doing the folding and generation itself internally. It's worth sitting with alongside the "grad student project" critique below. Whether you read this campaign as evidence of Claude's own scientific capability, or as evidence that a capable general model can now direct an existing open-source toolchain competently and at scale, depends on which of those two things you think is the more interesting claim. Both readings are consistent with the same data.
ML researcher Ravid Shwartz-Ziv sharpened that same orchestration point into a more pointed critique on X, reposted by Hugging Face's Lewis Tunstall. His argument: the open-source generators Claude directed "did most of the lifting," and several of them already report hit rates in a similar range on their own, before Claude was involved at all. He names PXDesign, RFdiffusion, Genie, and BoltzGen specifically, attributing their origins to labs including the Baker lab, Columbia, MIT, and ByteDance Seed. That last, specific claim — that the standalone tools already publish comparable hit-rate numbers — remains unverified, so treat it as an assertion from a credible-seeming source rather than a confirmed fact. His sharper, more checkable point is architectural. Since these generators are all trained on PDB-derived structural data, ensembling several of them doesn't necessarily diversify away a shared blind spot: "they fail together" on the same kind of target. That lines up with the pattern already visible in this campaign's own results, where the targets that worked were the well-studied ones. His proposed restatement of the valid claim: "an agent can now drive this stack competently in the regime where the stack already works." That's a narrower claim than "Claude designs proteins," and one this post's own numbers don't obviously contradict.
A pushback worth weighing: how novel were the targets?
A verified X account under the name Melinda B. Chu, describing herself as a researcher, posted a pointed critique of the announcement. Her reading: this is "a grad student or Post-doc project... first-year grad student or college student project," and the targets were "primarily known targets AND Adaptyv Bio already knew how to synthesize them b/c they were known targets." She asked why the result is being treated as a major accomplishment rather than routine work. Her professional credentials and affiliation are not independently confirmed. The factual part of her claim, though, is checkable against Anthropic's own text, and it's worth separating what's verifiable from what's a judgment call.
On target selection, Anthropic's own methodology confirms the core of her point: "We began our protein design campaign by selecting multiple targets that are commonly used in protein design benchmarks, including all of Adaptyv Bio's BenchBB. Because these targets have been studied extensively, we can compare our results against published hit rates and affinities." Of the roughly 15-16 targets used, Anthropic names only 15-PGDH and GDF-8 as targets chosen specifically for not having established prior art, "to ensure Claude was able to design against targets without drawing upon pre-recorded successes in its training data." So yes — most targets were established benchmarks. Anthropic says so directly, framing it as a deliberate choice for comparability rather than something concealed. Her inference that Adaptyv Bio already had synthesis and assay protocols worked out for these targets is plausible, since several of them are targets Adaptyv had already run public competitions on. But that specific claim isn't stated outright in Anthropic's materials, so it's a reasonable extrapolation rather than a confirmed fact.
In a follow-up post, Chu sharpened the point: "This is repeating an open contest from a few months ago. The person that won that contest didn't say they were curing all diseases." That specific claim — that RBX1 was already the subject of a public contest Anthropic's campaign then re-ran — is, again, confirmed directly by the announcement: the RBX1 comparison is against that same Adaptyv Bio competition. Her broader point is a mismatch in scale between the claim and the evidence. Anthropic's own public messaging, including Dario Amodei's stated ambition to cure most human disease within roughly 5-10 years, is far grander than "matched a known benchmark's winning entry on one target out of fifteen." That gap between a modest, specific experimental result and the scale of the ambition it gets cited in support of is a fair thing to hold in mind. A narrow win on a repeated benchmark is real evidence of something, but it isn't itself evidence of curing disease. Nothing in the announcement claims otherwise directly; the connection is made by the surrounding narrative, not by this specific result.
Where her framing is a judgment call rather than a checkable claim is the "grad student project" characterization itself — that's an assessment of significance and difficulty, not a fact this post can verify or refute. What the campaign's own numbers offer as context: on RBX1, one of the targets in question, Claude's design beat the winning entry among 245 submissions to a public competition. Those submissions presumably included exactly the population of researchers, grad students and postdocs among them, that her comment invokes as the appropriate comparison. Using well-studied targets does make the comparison more legible, since there's a known baseline to beat. It also means this campaign is not evidence of Claude succeeding on genuinely novel biology — a narrower and more accurate claim than "Claude accelerates protein design" might suggest to a casual reader. Both things can be true at once.
The RBX1 result is the specific one to remember
Against RBX1, a protein involved in targeted protein degradation, Mythos Preview in single-target mode hit 40% — and Adaptyv Bio had run a public competition on the same target where human entrants achieved a 3.7% hit rate across 245 submitted designs. Claude's top-ranked RBX1 design outperformed that competition's winning entry on affinity. That's a specific, falsifiable claim against a specific, dated public leaderboard — the kind of comparison this blog has repeatedly asked for and rarely gotten when a lab claims to beat "the field."
Where the two models disagreed, and Anthropic doesn't know why
The most candid line in the report concerns TNFα, a clinically important, structurally difficult target (it's the basis for Humira). Opus 4.8 designed working binders — including some that bind human, cynomolgus monkey, and mouse TNFα simultaneously, useful for animal studies — while Mythos Preview, the model that outperformed Opus 4.8 everywhere else in this campaign, failed on it entirely. Anthropic's own sentence: "We're not sure why Opus 4.8 was successful on this target and Mythos Preview was not." That's a real admission of a capability being non-monotonic across their own model line on a specific task, stated plainly rather than smoothed into a composite "our models are strong at protein design" claim. It's also a useful caution against reading any single hit-rate number as a stable ranking: on this one target, the "worse" model won outright.
Claude also produced 15 confirmed binders using β-sheet structures across six targets — a harder, misfolding-prone motif most computational design defaults to avoiding in favor of α-helix bundles — and the report is equally direct about failure: zero of 90 designs against maltose-binding protein were confirmed to bind (one showed a weak, reproducible signal), and its BBF-14 binders, while real, only reached modest sub-micromolar affinities against a target specifically chosen for being adversarially hard to design against.
The chemistry task: no autonomy claim, just a stopwatch
The second experiment used Opus 5, the generally available model — not a gated research variant — and tested something narrower: given a contract lab's raw NMR and LC-MS instrument files and a two-sentence prompt, could it reproduce the lab's own analysis? It returned results in 23 and 19 minutes, with hydrogen counts within 0.08 ppm of the lab's figures and a purity reading of 96.4% against the lab's 96.33%. Two details make this more than a speed claim: Claude independently proposed the same heavy-water follow-up experiment the lab had run three days after its own first measurement, and when it ran that follow-up data, it caught and corrected its own earlier overstatement — its first pass claimed four flagged peaks had all disappeared; its self-check showed only two had. That's a specific, visible instance of self-verification catching a real error, not just a faster final answer.
The data release backs the claim up further than the blog post does
Anthropic followed the announcement with a full data release — 1,440 designs in total (900 from Mythos Preview, 540 from Opus 4.8) across 16 targets, one more than the announcement's headline 15. The extra target, mature GDF-8, is included with its designs and provenance, but its wet-lab measurements are explicitly excluded because the antigen aggregated and bound assay surfaces non-specifically. That's an inconclusive result reported as inconclusive rather than dropped without comment. The two-CRO verification is also more rigorous than "two labs agreed." Adaptyv Bio and Twist Bioscience ran genuinely different assay methodologies on the same designs: cell-free expression with the design itself immobilized for SPR/BLI kinetics at Adaptyv, Fc-fusion expression with a six-point antigen titration under capture SPR at Twist. A binder counted as confirmed had to clear two different experimental setups, not just two different buildings running the same protocol.
What's released goes well past the numbers in the announcement: raw sensorgrams and report images from both vendors per design, the final per-design comparison and verdict between the two CROs, co-folded structure predictions from ten different structure predictors at five seeds each (113,550 predicted structures with full PAE matrices), and step-level design provenance. Genuinely unusual, it also includes the exact campaign prompts, the kickoff messages, and the full external resource corpus Claude had access to during the campaign — cited papers, ProteinBase collections, method papers, each with its own redistribution and licence notes. Data and documentation are CC BY 4.0. That's a reproducibility package closer to what an independent academic group would need to actually re-run the analysis than what usually accompanies a lab's capability announcement. It strengthens the credibility case in the same direction as the two-vendor validation itself: nothing here has to be taken on Anthropic's word alone.
The dual-use gate, stated plainly
Anthropic is explicit that protein design and "other dual-use research biology capabilities" remain withheld from Claude Fable 5's general access, available only through a "trusted access program" it says is coming soon. That's the same shape of gating this month's cyber-capability coverage has tracked repeatedly, applied here to biology instead of cyber, and for the same stated reason: autonomous research capability is dual-use. It's also why the wet-lab campaign described above ran on Mythos Preview and Opus 4.8 — the most capable models are still blocked from this class of task pending that access program.
What to expect next
- Watch for the promised follow-up characterization. Anthropic states outright it intends further work to confirm these hit rates and affinities — a real commitment to re-checking its own headline numbers rather than letting a first result stand as final.
- Watch whether the TNFα asymmetry recurs on other targets. One unexplained model-capability inversion is a data point; a second would suggest something structural about how these models approach binder design differently, not noise.
- Watch the trusted-access program's actual terms, since that's what determines whether outside researchers — not just Anthropic's own team — can independently reproduce any of this on a model more capable than the ones used here.
- Watch for peer review. Nothing here has gone through a journal or formal external review yet; the technical reports and open data make that step possible, but publication and independent critique remain the actual test of whether these results hold up under scrutiny.