Claude enters protein design and delivers an impressive debut performance.
Just now, Company A released a new blog post, providing a detailed retrospective on their joint journey into the AI4S space with Mythos and Opus.

Quickly flip to the results page, and indeed, there's something there:
Claude designed 1,320 protein candidates from scratch, which were then sent to external laboratories for production and validation.
The results showed 354 successful target combinations, with 14 out of 15 targets achieved, reaching a maximum hit rate of 35.1%.
The typical hit rate for current protein design projects is usually only 10% to 15%.
Even more impressively, after designing the new molecule, it hasn’t even thought about clocking out.
Faced with the NMR and LC-MS raw files provided by the lab, Claude processed them in 23 minutes and 19 minutes respectively, after reading only two sentences of instructions.
The final measured purity of the sample is 96.4%, while the laboratory's reported result is 96.33%.
Just 0.07 percentage points away.
In other words, Claude's results closely match the laboratory results.

It's not about prediction—it's about designing proteins from scratch.
When people think of protein, many immediately think of Google DeepMind’s AlphaFold project.
To clarify, AlphaFold predicts protein structures, primarily addressing the question: given a sequence of amino acids, what structure will it fold into?
This time, Claude was given a completely different task:
The target protein is placed here; design a completely new protein from scratch that can bind to it precisely.
Claude is designing something called a minibinder, which is a miniature binding protein.
According to publicly available information, minibinders are artificially designed small proteins, typically consisting of only 50 to 70 amino acids, with a compact structure and a single function.
Its task is straightforward: to bind with extremely high precision and strength to a specific site on the target protein.
Don’t underestimate this “binding” action—many drugs work precisely because of this molecular-level specificity:
Stick to bad proteins to inactivate them, or bind to good proteins to activate them, or directly deliver drug molecules to the site of disease for release.
The problem is that this task is quite taxing.
Protein designers must not only orchestrate a series of specialized models but also repeatedly generate, optimize, and screen candidates—a process that demands significant expertise.
Each target requires experts to spend weeks or even months sifting through a large number of candidate molecules to identify just a few that are truly effective.

Anthropic has now handed Claude the entire toolkit.
The participants in the experiment were Mythos Preview and Opus 4.8, which had access to academic papers, online resources, and multiple specialized protein models, along with substantial GPU computing power.
After humans provided the initial task, Claude took over on its own. The overall result was:
Claude successfully targeted 14 out of 15 targets, yielding at least six high-affinity binding proteins and matching or surpassing previous best results in at least four cases.
There are two specific methods:
Multi-target mode: Enable Claude to handle multiple objectives simultaneously within the same 48-hour task, with up to 12,500 H100 GPU hours available.
Single-target mode: Focus on one target at a time; all tasks run in parallel for 24 hours, with a maximum of 2,500 H100 GPU hours per target.
Experiments have shown that focus is indeed better than multitasking.
Under the multi-target mode, the overall hit rates for Mythos Preview and Opus 4.8 are 26.7% and 22.6%, respectively.
After switching to single-target mode, Mythos Preview surged to 35.1%.

Performance on certain benchmarks has even surpassed previous human achievements.
For example, RBX1 is a target involved in regulating protein degradation. In a design competition hosted by Adaptyv Bio, the overall hit rate among all participants was only 3.7%, whereas Mythos Preview achieved a 40% hit rate when challenging it independently.
Not only is it easier to hit, but Claude's top-ranked design is a high-affinity binding protein with binding strength that surpasses the winner selected from 245 competing designs.

Even more challenging is TNFα, an inflammatory signaling protein released by the immune system.
Due to the suitable binding site being hidden in the groove formed by the two proteins, multiple expert teams have previously encountered obstacles here.
This question, Mythos Preview, was not solved, but was instead conquered by Opus 4.8:
Opus 4.8 successfully designed multiple effective proteins, some of which can simultaneously bind to human, cynomolgus monkey, and mouse TNFα.
This cross-species compatibility is crucial because it eliminates the need to redesign the approach each time you switch species for animal testing.
But the question then arises:
Why did the overall stronger Mythos Preview fail, while Opus 4.8 succeeded?
Company A admitted that they don't fully understand it either.

It can only be said that the research capabilities of large models are still uneven.
As Claude continued to tackle more challenging β-sheet designs, it successfully created 15 effective binding proteins across six targets.
But when switching to maltose-binding protein (MBP), Claude completely failed—
Ninety designs had no successful confirmations; only one showed a faint signal.
So, it's still too early to say "input a target, and Claude will automatically generate a new drug," but Claude has already demonstrated:
It can autonomously orchestrate expert models, completing design workflows that previously required experts to spend significant time organizing and filtering.
This step is already significant.
After designing the new molecule, Claude began reading the experimental data.
However, Company A's experiment did not stop there.
In addition to testing whether Claude can design new things, they also wanted to see:
Given a synthesized compound, can Claude independently read experimental data to determine what it is and how pure it is?
This time, the universally available Claude Opus 5 is here.
The task comes from two very routine but tedious tasks in analytical chemistry:
NMR spectroscopy: Observe the signals of hydrogen atoms in the molecule to confirm whether the synthesized compound is the target molecule.
Liquid chromatography–mass spectrometry (LC-MS): Separates the different components in a sample, then measures their molecular weights and concentrations to determine whether the sample is pure and what else it contains.
This task requires chemists to manually process raw data in proprietary formats, step by step calibrating, identifying peaks, integrating, verifying, and finally compiling a report.
It involves multiple steps and requires significant experience; even a slight deviation can lead to incorrect conclusions.
Anthropic provided only the raw files returned by the Claude lab and two sentences of task instructions.
No vendor software is used, and no operator is present to provide guidance (P.S. Relevant data and prompts are open-source).

Claude completed the NMR and LC-MS analyses in 23 and 19 minutes, respectively.
For NMR, Claude identified 18 signal peaks and calculated the number of hydrogen atoms corresponding to each peak, with results deviating by no more than 0.08 hydrogens from the laboratory data.
It also identified that four of the broad peaks may originate from hydrogens bonded to nitrogen or oxygen, and suggested adding deuterium oxide for further verification.
This is a common process of elimination: after adding heavy water, the signals of certain hydrogen atoms weaken or disappear, allowing chemists to deduce the molecular structure.
Interestingly, the lab later conducted the same verification.
After verification, Claude initially thought all four signals had disappeared, but upon self-checking, realized this was incorrect and corrected it to only two signals having disappeared, ultimately reaching the same conclusion as the laboratory.

LC-MS is more straightforward here.
For vendor files without publicly documented formats, Claude first determined how the data was encoded, then verified the file was read correctly by reproducing all 2,664 scan records.
It then provided the chromatogram, mass spectrum, purity table, and molecular weight, along with a reusable analysis script.
Ultimately, Claude measured the sample purity at 96.4%, while the laboratory result was 96.33%.
Just 0.07 percentage points away.

Company A concluded that:
Both experiments point to one thing: AI is lowering the professional barriers, costs, and time required for life science research. In chemical analysis, Claude is taking over tasks previously dependent on manual labor; in protein design, Claude can achieve end-to-end ligand design with minimal input, with some results matching or even surpassing previous best designs.
Although we are still far from actual new drug development, AI has at least shown us the possibility of accelerating this process.
Left a little Easter egg too.
At the end of the blog, Company A also highlighted a key point for scientists.
This protein experiment used Mythos Preview and Opus 4.8; Fable 5, which has greater capabilities, is currently not available to general users for life science tasks.
The reason is straightforward: the ability to design drug proteins could also be used for dangerous biological research.
However, Company A has announced that it is preparing a dedicated access program for scientists, with more details to be revealed later.
Full technical report: https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf
Related prompts and open-source data address: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/tree/main
Reference link:
[1] https://www.anthropic.com/research/Claude-accelerates-protein-design
[2]https://x.com/AnthropicAI/status/2089842387845804246
This article is from the WeChat public account "Quantum Bit," authored by Yi Shui.
