Artificial intelligence is moving into an area that once required highly specialized scientists, complex software and extensive laboratory work.
Anthropic has now tested whether its Claude AI models can design new proteins that bind to specific biological targets. The results are notable: Claude produced confirmed protein binders for 14 of 15 targets that could be evaluated, with independent laboratories physically producing and testing the AI-designed proteins.
The experiment does not mean Claude has created new medicines. However, it shows how AI could speed up one of the early and difficult stages of drug discovery.
What Did Claude Actually Do?
The experiment focused on designing protein binders, also known as minibinders.
A protein binder is a small protein designed to attach to another specific protein. In medicine and biotechnology, this can be useful because many diseases involve proteins that scientists want to block, activate or otherwise control.
Finding a protein that binds strongly to the right target can be difficult. Scientists traditionally have to generate many candidates, predict their structures, test them computationally and then send promising designs to a laboratory.
Anthropic wanted to see whether Claude could manage much of this computational process itself.
The company used two models for the main protein-design campaign: Claude Mythos Preview and Claude Opus 4.8.
Claude Designed 1,440 Protein Candidates
The public dataset released by Anthropic contains 1,440 de novo miniprotein designs, ranging from 50 to 120 amino acids in length.
The designs targeted 16 different biological targets.
However, the laboratory results for 120 designs targeting mature GDF-8 were inconclusive because the target protein aggregated and showed nonspecific binding in the testing system.
That left 1,320 designs across 15 targets for the main evaluation.
The results were then independently assessed using laboratory experiments performed by Adaptyv Bio and Twist Bioscience.
354 AI-Designed Proteins Actually Bound Their Targets
This is the part of the experiment that attracted the most attention.
Of the 1,320 designs tested across the 15 evaluable targets, 354 were confirmed as binders.
Claude produced at least one successful binder against 14 of the 15 targets.
That works out to an overall hit rate of approximately 26.8%.
Anthropic compared this with typical protein-design campaigns, which it says often achieve hit rates of roughly 10% to 15%.
The comparison should not be interpreted as a universal industry benchmark because protein-design success rates can vary depending on the target, assay and experimental setup. Still, the result suggests that Claude was able to generate a relatively high proportion of candidates that worked in laboratory tests.
The Proteins Were Not Just Computer Predictions
One of the most important details is that the researchers did not stop at computer simulations.
Adaptyv Bio describes its role as taking the digital protein sequences through an automated wet-lab process, including DNA production, protein expression, measurement and quality control.
Twist Bioscience independently evaluated 1,260 protein binders across 15 targets in less than three weeks.
This matters because a protein that looks promising in a computer model can fail when scientists actually try to produce it.
The laboratory testing therefore provided a much stronger test of Claude’s designs.
How Did Claude Design the Proteins?
Claude was not given a blank screen and simply asked to invent proteins.
The experiment started with a detailed expert-written protocol. Anthropic says the prompt was approximately 30,000 tokens long.
Claude then had access to scientific literature, computational resources and specialist protein-design tools.
The AI could research the targets, identify possible binding regions, generate candidate structures and sequences, run computational analyses, optimize candidates and rank the designs for laboratory testing.
This distinction is important.
The breakthrough was not that Claude replaced every specialized protein-design model.
Instead, Claude acted more like a scientific project manager and researcher, coordinating several existing tools to solve a larger problem.
Anthropic’s approach is consistent with its broader Claude Science strategy, where an AI assistant can connect to scientific databases and specialist tools and divide complex research tasks among different AI workers.
Claude’s Success Rate Varied by Experiment
Claude did not perform equally well in every setup.
When Mythos Preview worked on multiple targets at the same time, it achieved a 26.7% hit rate.
Opus 4.8 achieved a 22.6% hit rate in the multi-target campaign.
When Mythos Preview was instead focused on individual targets, its overall hit rate increased to 35.1%.
This suggests that giving an AI system a narrower objective can sometimes allow it to spend more computational effort exploring and improving designs for that specific target.
One Target Produced a 40% Hit Rate
One of the strongest individual results involved RBX1.
In the focused experiment, Mythos Preview achieved a hit rate of about 40%.
Adaptyv Bio notes that this compares with a much lower result from participants in an earlier protein-design competition involving the same target.
Anthropic also reported that some of its strongest designs had binding affinities that compared favorably with previously published results.
These results are important because the experiment was not simply about generating any protein that could attach to a target.
The strength of the interaction also mattered.
Claude Also Produced Different Types of Protein Structures
Another interesting finding was the structural variety of the designs.
Protein design systems often favor certain structural patterns because they are easier to generate and predict.
Anthropic reported confirmed binders containing significant amounts of beta-sheet structure, including designs against several different targets.
This suggests that Claude was able to explore more than one structural solution instead of repeatedly producing similar protein shapes.
Claude Also Designed a Binder for TNF-Alpha
One of the more interesting biological targets was TNF-alpha, a protein involved in inflammation.
TNF-related biology is already an important area in medicine, with existing treatments targeting this pathway.
Opus 4.8 produced binders against TNF-alpha, and some of the designs were able to bind versions of the protein from humans, mice and cynomolgus monkeys.
That could potentially be useful in future research because researchers often need models that allow them to study candidates across different species.
Interestingly, Mythos Preview did not succeed against the same target, despite performing better overall in the broader campaign. Anthropic said it does not yet know why the two models behaved differently on this target.
Claude Did Not Succeed Against Every Target
The failures are just as important as the successes.
For example, Claude did not produce a confirmed binder against maltose-binding protein (MBP) in the main evaluation.
This shows that the system is not capable of reliably solving every protein-design problem.
The mature GDF-8 experiment also could not be properly evaluated because the target aggregated and produced nonspecific interactions during testing.
These limitations are important because headlines such as “Claude can design any protein” would be misleading.
The experiment demonstrated strong performance across a set of targets, not universal protein-design capability.
Why Is This Important for Drug Discovery?
The early stages of drug discovery can involve a huge amount of trial and error.
Scientists need to identify biological targets and then find molecules or proteins capable of interacting with those targets in useful ways.
AI could reduce the amount of time needed to explore this enormous search space.
Instead of manually designing candidates one by one, an AI agent could potentially run a cycle like:
Research → Design → Predict → Optimize → Rank → Test
The computer handles much of the repetitive computational work.
Scientists can then spend more time on laboratory validation, experimental decisions and interpreting the results.
This is why the Claude experiment could be important even though none of the proteins is currently a medicine.
It demonstrates a potential way to make the early discovery process faster.
But These Are Not New Drugs
This is perhaps the most important limitation to understand.
A protein binder is not the same thing as a drug.
The fact that one of Claude’s proteins binds successfully to a target does not prove that it can treat a disease.
A promising protein would still need extensive testing to determine whether it:
- works in cells
- produces the desired biological effect
- remains stable
- reaches the right part of the body
- avoids unwanted immune reactions
- has acceptable toxicity
- works in animal studies
- is safe enough to enter human clinical trials
Anthropic itself describes protein binding as an early step rather than the end of drug development. Independent coverage has also emphasized that the experiment does not demonstrate autonomous drug discovery.
So the 354 successful binders should be viewed as potential starting points for further research, not 354 new medicines.
Claude Also Analyzed Laboratory Data
The protein-design experiment was not the only scientific test Anthropic announced.
The company also tested Claude Opus 5 on analytical chemistry.
The model was given raw NMR and LC-MS data and asked to analyze it.
Claude reportedly completed the NMR analysis in about 23 minutes and the LC-MS analysis in about 19 minutes.
Its calculated purity for one sample was 96.4%, compared with 96.33% in the contract laboratory’s analysis.
This illustrates another potential use of AI in science.
AI does not necessarily have to discover a new molecule. It can also help researchers process, interpret and validate the large amounts of data produced by laboratory instruments.
Anthropic Released the Data
Another notable part of the announcement is that Anthropic did not keep all of the protein-design results private.
The company released a public dataset containing the designs and experimental information.
The dataset includes:
- protein sequences
- binding results
- binding kinetics
- raw experimental data
- structure predictions
- design provenance
- target information
- laboratory measurements
The release contains 129,003 files and about 9.9 GB of data in its unpacked form, according to the dataset documentation. It is released under a CC BY 4.0 license.
This makes it possible for researchers and other scientists to examine the results themselves.
The Biggest Change May Be the Role of the AI
The most interesting part of this story is not necessarily the protein sequences.
It is the changing role of AI.
Traditional AI tools are often designed to perform one specific task.
One system predicts protein structures.
Another generates sequences.
Another evaluates binding.
Another analyzes laboratory measurements.
Claude’s role in this experiment was different.
It could coordinate the tools and decide what to do next.
That makes the AI less like a calculator and more like a junior research scientist working with a large collection of specialized software.
This is also where the broader Claude Science strategy becomes relevant. Anthropic has been building workflows where Claude can connect scientific databases, use specialist tools and coordinate multi-step research projects.
What Could Happen Next?
The next challenge is moving beyond binding.
Future AI systems could potentially design proteins while optimizing several properties at the same time.
For example, researchers might want a protein that:
- binds strongly to a target
- avoids similar proteins
- remains stable
- can be produced efficiently
- works under biological conditions
- produces a specific biological effect
That is much harder than simply designing something that sticks to a target.
If AI systems become good at handling these additional constraints, they could become increasingly useful in pharmaceutical research.
But every computational prediction will still need experimental validation.
AI Biology Also Creates Safety Questions
There is another side to this progress.
Protein-design capabilities can be used for beneficial research, but advanced biological design systems can also have dual-use risks.
The same ability to design a protein that interacts with a biological target could potentially be useful in areas that require much stricter safeguards.
Anthropic has therefore placed restrictions around some advanced biological work and has discussed controlled access to its most capable scientific systems.
This creates a difficult balance.
Scientists want AI systems that are powerful enough to accelerate research, but those same capabilities need safeguards so they are not easily misused.
What This Experiment Really Proves
The safest way to interpret Anthropic’s announcement is not:
“Claude invented new drugs.”
It is:
Claude demonstrated that a general-purpose AI system can coordinate a sophisticated protein-design workflow and produce many candidates that work when physically tested.
The numbers are impressive.
Out of 1,320 experimentally assessed designs, 354 were confirmed binders, and at least one successful binder was found for 14 of 15 evaluable targets.
But the experiment also had failures, and none of the results means a new medicine is ready for patients.
The bigger significance is what happens when AI moves from simply answering scientific questions to actually running parts of the scientific discovery process.
If these systems continue improving, the future of drug discovery may involve scientists working alongside AI agents that can research a target, design hundreds of candidates, analyze experimental data and suggest what should be tested next.
Claude’s protein experiment is an early demonstration of what that future could look like.
