Nobel Laureate Bets $94.6 Million on AI-Driven Protein Design Beyond Nature

icon MarsBit
Share
AI summary iconSummary
AI and crypto news broke as AI BioDesign, led by 2024 Nobel Chemistry winner David Baker and Jay Shendure, launched in Seattle with $94.6 million. The institute uses AI and high-throughput biology to design proteins that nature never evolved. Instead of scaling up AI models, the team conducts parallel experiments in labs, processing millions of DNA sequences. The project employs a closed-loop system of AI design, laboratory testing, and model learning. On-chain observers may track how this data-driven approach bridges AI speed with real-world validation.

AlphaFold has predicted the structures of nearly 200 million proteins found in nature.

David Baker wants to take a different path: instead of predicting what nature has created, he aims to create things nature has never made.

On a morning in August this year, research assistant Jack Boylan pulled up the previous night’s experimental results.

Next to the DNA sequencer in the lab, test tubes contain a large number of artificially designed DNA sequences that have never evolved in nature. Which ones work and which are useless is determined entirely by this machine, which reads and reports them all in one go.

In a recent experiment, they pooled 6 million of these design sequences into a single test tube and sequenced them all at once. Using traditional methods, sequencing just this batch would have taken years.

The organization where Boylan is located is called AI BioDesign, which debuted in Seattle on September 3, just minutes away from the Allen Institute headquarters.

Behind it is a $94.6 million investment from the Paul G. Allen Estate, led by 2024 Nobel Prize in Chemistry winner Baker and genomicist Jay Shendure.

Synthetic biology

Left: David Baker, 2024 Nobel Prize winner in Chemistry. Right: Jay Shendure, genomicist. Both co-lead AI BioDesign.

In an era where computing power is everything, they don’t plan to spend nearly $100 million on training a larger model—they intend to focus their efforts on that test tube.

Hijack cellular machinery and perform high-concurrency operations in a test tube.

Their strategy is to create a runaway feedback loop:

The AI proposes a design, the lab fabricates and measures it, the results are fed back to the model, and the model then decides what to fabricate next.

It sounds like the familiar "Design-Build-Measure-Learn" cycle, but what truly sets it apart is how this loop actually spins.

There are no automated workshops filled with robotic arms. “Scale doesn’t come from robots—it comes from doing things in parallel inside test tubes,” says Executive Director Jesse Gray.

所谓的 parallelization within the tube, in Shendure’s words, is hijacking the cellular assembly line that evolution has left to humanity and using it to build and test millions of molecules.

The process involves dumping millions of design sequences into a single test tube, allowing cells to produce each one individually, then running them all under the same conditions. Finally, a sequencer reads out which sequences are functional and which are defective.

Thus, when the model learns the "designed rules," it is presented with vast numbers of examples rather than the small, naturally occurring set.

What do these designed molecules look like?

The image below shows a designed fusion protein from Baker Lab: the purple and pink portions were designed by AI, and the white portion is amylin completely enclosed by it.

Synthetic biology

AI-designed surface rendering of an amylin-binding protein.

Experiments no longer verify conclusions; they merely feed the model.

More interestingly, this production line's KPI.

Laboratory teams are organized into small groups of five or six members, each tackling a specific biological design challenge, with a machine learning team on standby to provide immediate support.

How is the performance of each round calculated? It’s not just about how many usable molecules are produced, but also about how much the model learns from that round.

The traditional laboratory approach follows the logic of forming a hypothesis first, then designing experiments to test it.

AI BioDesign reversed the order: experiments begin as "practice problems" for the model, cycling repeatedly through the "design-build-measure-learn" loop.

At the end of each round, the machine learning team identifies the sequences where the model was least confident and most likely to learn something new, and selects them for the next round of testing. The tested data is fed back into the model, which is then updated to generate the next batch of designs.

The goal is to maximize the amount of information gained from each experiment, rather than quickly putting together a finished product.

In the past, humans directed the lab bench; now, the lab bench follows AI's instructions.

Synthetic biology

Research assistant Jack Boylan (left) from the Allen Institute and Executive Director of AI BioDesign Jesse Gray, in front of a DNA sequencer in the lab.

He created proteins not found in nature as early as 2003.

This is not a story about "AI creating proteins that do not exist in nature."

Baker accomplished this 23 years ago.

In 2003, his team used the Rosetta software to design a protein called Top7: 93 amino acids long, with a folding pattern never before seen in any known protein.

Half of the 2024 Nobel Prize in Chemistry was awarded to Baker for computational protein design, and the other half to Demis Hassabis and John Jumper of the AlphaFold team for protein structure prediction.

Synthetic biology

On October 9, 2024, David Baker received a call from the Nobel Committee.

If AlphaFold addresses the question of "reading"—predicting what shape a sequence of amino acids will fold into—then Baker addresses the question of "writing"—what shape should I design to achieve a specific function?

In recent years, that half of the work has been elevated to godlike status; AlphaFold2 and other large models from the same period followed the same approach: they relied entirely on data already accumulated by humans.

You can't get away with just "writing" it. If it doesn't exist in nature, you won't find it in any database—you have to create it yourself and test it yourself.

Thus, it has remained stuck in a quagmire of reality:

AI can generate hundreds of thousands of candidate designs in a single night, while a traditional lab considers hundreds of designs in a year to be highly productive.

The vast gap between design and verification is the ceiling this industry has consistently struggled to break through.

In Baker’s words, the real breakthrough of this project is that AI’s speed has for the first time begun to match the experimental capabilities of synthetic biology.

On one side, millions of candidate designs are generated each day; on the other, millions are tested in a single tube—these two gears have finally meshed.

We do not chase general-purpose large models; we heavily invest in experimental infrastructure.

While everyone is chasing a general large model that can answer any question—and even simulate an entire virtual cell—Baker and AI BioDesign chose a different path:

Create a set of small models, each focused on specific hard problems, and generate vast amounts of new data tailored to each problem to train them.

The concept is not hard to understand.

The samples left by natural evolution are merely a tiny fraction of solutions it discovered over billions of years; trying to derive universal design principles from this small subset inherently leaves significant gaps.

To fill this gap, you need to create the data you're missing—and creating data doesn't rely on stacking GPUs, but on test tubes, sequencers, and lab benches.

In the previous stage, everyone competed for computing power infrastructure; in the next stage, what we need to compete for is experimental infrastructure.

More counterintuitively, even if you win this round, all the results are given away for free. The models, datasets, experimental methods, reagents, and benchmarks are all open.

What if someone uses this data to create a profitable product?

President Rui Costa said casually, "If many companies use this data to make the world a better place, that’s our luck."

Even amid rapid advancement, there are hidden concerns. Mass-producing molecules that do not exist in nature must guard against unforeseen ecological impacts and malicious misuse.

On this chain, there’s an unavoidable step: synthesizing DNA. No matter how many designs you create, someone ultimately has to turn them into physical reality.

What’s in place here is the industry’s screening system. When you place an order to synthesize a segment of DNA, the system compares it against known toxin and pathogen genes—and blocks it if there’s a match.

A study published in Science in October 2025 showed that the Microsoft team used publicly available protein design tools to rewrite 72 target proteins into approximately 72,000 synthetic homologous sequences, preserving their structure and function while drastically altering their sequences, causing four screening software tools to miss the majority.

The team then collaborated with four synthetic companies to patch the issue, resulting in a significant increase in detection rates.

Synthetic biology

Microsoft team's AI protein design red teaming process: generate synthetic homologous sequences from target proteins, then submit them to screening software for validation.

The problem is that patches can only address vulnerabilities that have already been discovered. Design capabilities are advancing exponentially, while defenders are always playing catch-up.

Baker's claim is: All synthetic DNA should be registered and documented, recording both the sequence and the purchaser.

He believes this is a practical barrier against abuse.

Open the drawer of biology 18 months later

The bet has been placed.

Costa set a deadline of 18 to 24 months to deliver tangible progress and to open access to researchers worldwide within five years.

Baker’s vision list includes: therapies for new diseases developed in weeks rather than decades, plastics in the ocean broken down by enzymes, and bio-computers that consume far less power than silicon chips.

These are still just directions at this point. Baker himself acknowledges that along this path, a batch of molecules will emerge that work but cannot yet explain why they work.

Can be made, but hard to explain.

Darwin once marveled at nature's diversity with the phrase "endless forms most beautiful."

But in Baker’s view, those most beautiful forms are just a tiny fraction that have been tested over billions of years—most of the drawers remain unopened.

Now, with the help of AI, he has opened the drawer, flipping through six million entries at once.

AlphaFold has predicted the structures of over 200 million proteins—this path to “understanding nature” has reached its end.

Baker believes the next major breakthrough won’t come from models, but from data: a production line that continuously generates new data.

Reference materials:

https://www.wired.com/story/nobel-prize-protein-design-now-using-ai-to-create-molecules-beyond-nature/?utm_source=chatgpt.com

https://alleninstitute.org/news/ai-biodesign-accelerator-combines-experimental-biology-and-artificial-intelligence-to-learn-natures-design-rules?utm_source=chatgpt.com

https://www.nobelprize.org/prizes/chemistry/2024/press-release/b/?utm_source=chatgpt.com

This article is from the WeChat public account "New Intelligence Yuan" (ID: AI_era), authored by Yuan Yu.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.