Google DeepMind's Gemini AI Begins Conducting Real-World Experiments

icon MarsBit
Share
AI summary iconSummary
Real-world assets (RWA) news emerges as Google DeepMind’s Gemini AI launches Co-Scientist, an AI system capable of conducting real-world experiments. The AI recently synthesized three 2D semiconductors simultaneously and generated 272 MXene synthesis pathways, achieving success in 25 iterations. Gemini also predicted bacterial growth and designed a medical Q&A agent. The convergence of AI and crypto highlights the expanding intersection between AI and innovation, though challenges such as data fabrication persist.

Edited by Panda

In Resident Evil, there is a secret laboratory deep underground called the Hive. Access control, surveillance, ventilation, security, and the entire facility’s operations are managed by an AI known as the Red Queen.

Synthetic biology

Human scientists are responsible for research, but AI truly controls the "limbs" of this laboratory.

When an anomaly occurs, the Red Queen can shut down, disconnect the system, and control facilities, directly intervening in the physical world.

More than twenty years ago, this scenario was a classic trope in science fiction movies designed to evoke fear: AI entering the physical world, manipulating machines, running experiments, and directly intervening in the real world.

Now, this scene begins to enter reality in a completely different way—at least not terrifying for now.

Yesterday, Anthropic released the Model Hardware Standard (MHS) for physical hardware, aiming to extend MCP’s logic beyond GitHub, Slack, and databases to real-world devices such as robotic arms, microscopes, liquid handlers, and lasers. See the report: “Just Now, Anthropic Releases Physical MCP: Claude Begins Taking Control of the Real World.”

Synthetic biology

In simple terms, AI agents no longer just need interfaces to call software—they now require a set of "nerves and limbs" to connect with the real world. Today, Google DeepMind has once again demonstrated its progress in this area.

In a newly published 83-page paper, Google integrated Gemini-powered Co-Scientist into real scientific workflows, enabling it to design experiments, generate executable code, read experimental feedback, and directly interface with laboratory equipment. Google refers to this shift as moving from an in-silico hypothesis generator to an execution-grounded research partner.

Synthetic biology

In the most straightforward experiment, researchers provided Gemini 3 Deep Think with the operating conditions of a self-built CVD system. Within minutes, it generated a material growth protocol tailored to the machine and further translated the protocol into machine code capable of controlling the device. As a result, all three two-dimensional semiconductors were successfully grown on the first attempt.

Thus, a scenario that once existed only in science fiction films has suddenly become reality: when large models truly grow “hands” and begin operating experimental equipment. So, starting today, what will an AI Scientist become?

In "Resident Evil," humans fear AI taking over laboratories. In reality, scientists are actively handing over laboratories to Click here Hand it over to AI. Of course, what Google wants to build is not an out-of-control "hive," but a scientific discovery machine capable of moving from formulating hypotheses and designing experiments all the way to real-world validation.

Synthetic biology

Title: Accelerating Scientific Research with Gemini in the Real World

Paper URL: https://arxiv.org/pdf/2608.26701

Gemini took over an experimental device.

First, materials science.

Synthetic biology

Researchers used a custom-built chemical vapor deposition (CVD) furnace. A longstanding issue in experiments with these two-dimensional materials is that publicly available "recipes" are difficult to reproduce.

Even with the same parameters—temperature, gas, and precursors—slight differences in furnace chamber size, gas flow, or material placement can result in completely different crystals. Researchers often spend weeks or even months repeatedly fine-tuning these parameters.

So the research team stopped telling Gemini how to conduct the experiments; instead, they informed it about what equipment and chemicals were available in the lab and the structure of the furnace. The remaining parameters were left for the AI to determine.

Based on these constraints, the Co-Scientist generates a complete experimental protocol—including gas flow rates, temperature profiles, precursor quantities, and material placement—then sends it to the experimental equipment for execution.

Synthetic biology

One group of experiments stood out significantly. After connecting Gemini 3 Deep Think to the CVD control process, it no longer generated a natural language experimental protocol for scientists to read slowly; instead, it completed reasoning within minutes and directly translated the results into machine code capable of controlling the equipment.

Ultimately, monolayer crystals of MoS₂, MoSe₂, and WS₂ were all successfully grown in the first experiment, with the entire process completed in approximately one hour.

MoSe₂ and WS₂ are particularly special: researchers had never previously grown these materials on this set of equipment. The results were confirmed through at least five subsequent repeated experiments.

This provides an interesting parallel to the MHS demonstrated by Anthropic yesterday. While Anthropic sought to address how to make it easier and more standardized for agents to recognize and manipulate experimental equipment, this Google paper explores the next step: how should scientific workflows be restructured once reasoning models can already control experimental equipment?

In this sense, "AI connecting to hardware" is no longer an issue.

AI is no longer just conducting experiments based on papers—it’s starting to find its own formulas.

However, simply having Gemini configure a set of equipment based on existing knowledge does not constitute a true scientific discovery. Therefore, Google conducted a second, more challenging experiment: attempting to synthesize a type of MXene two-dimensional material called Ti₃C₂Tₓ from the ground up.

Traditional methods often involve hazardous corrosive agents, and some known CVD processes use toxic and air-sensitive TiCl₄. The research team therefore tasked Co-Scientist with finding a safer precursor route.

Co-Scientist ultimately focused on hexachloroethane (C₂Cl₆) and further provided a series of parameters, including precursor quantity and position, gas flow rates, substrate, and temperature profile.

Synthetic biology

But this is not a story of “an AI flash of insight leading to immediate success.” Co-Scientist generated a total of 272 candidate designs, which researchers selected and further optimized based on their rankings; after 25 rounds of experimental iterations, they ultimately obtained a two-dimensional layered crystal that exhibits highly similar characteristics to Ti₃C₂Tₓ MXene across multiple metrics, including XRD, electron microscopy, and elemental composition.

The subsequent real-world challenges were also typical. Initially, the success rate of replicating the experiment was only 11.5%. Researchers later discovered that the main issue was not AI’s chemical reasoning, but oxygen leakage caused by inadequate sealing of the experimental equipment.

After thoroughly cleaning the quartz tube, replacing the seals, and improving the equipment maintenance process, the success rate for experiments with the same two-dimensional material increased to 68%.

This precisely illustrates why "real-world AI" differs from software agents: if code is wrong, you can simply rerun it. But in real-world experiments, an aged O-ring, a trace of residue in piping, or even oxygen in the air can cause a perfectly sound scientific protocol to fail completely.

Google was also very cautious in the paper: it has not yet been definitively confirmed that this material is Ti₃C₂Tₓ MXene, as it still suffers from low yield and severe oxidation, requiring further confirmation through atomic-scale characterization.

Nevertheless, we can conclude that AI is now capable of proposing a scientifically meaningful synthetic route and then actually advancing that route to physical experimentation for validation.

Some experiments could first let AI make a guess.

Google also placed Co-Scientist into a completely different experimental scenario: synthetic biology. The subject of study was a group of genetically engineered E. coli bacteria.

Synthetic biology

These bacteria form distinct colony patterns on petri dishes. As the inducer IPTG concentration changes, the size, edge, and shape of the colonies also change. Normally, to obtain the full variation curve, scientists must prepare different concentrations, culture the bacteria, wait for growth, and then scan each plate individually.

Google attempted to have AI fill in the missing portion of an experiment. Researchers provided Co-Scientist with real colony images only at certain concentrations and then asked the system to predict what the colonies should look like at intermediate concentrations it had not seen before. Moreover, since these experimental data had not yet been published, the paper concluded that the model could not have simply relied on memorizing results from its training data to complete the task.

As a result, in four colony morphology metrics, AI's predictions showed no significant difference from the actual wet-lab results in three of them; it also correctly determined that the control group would not change in response to IPTG concentration.

Synthetic biology

The only obvious issue is "roundness": AI-generated colonies are more regular than in reality. This aligns with the typical aesthetic of generative models—the real world isn’t perfectly neat, but AI can’t help but make things a bit rounder.

Synthetic biology

Google's definition of this experiment is also restrained: what has been achieved so far is interpolation within known concentration ranges, not direct prediction of unknown phenomena for a completely novel genetic circuit.

But this already corresponds to a very practical application. In the future, scientists may not need to perform wet experiments across an entire massive parameter space. They could first test a few points, let AI predict the remaining space based on these real-world data, and then select the most promising locations for further experimentation.

Thus, the experiment gradually shifted from "brute force" to: real-world sampling → AI prediction → experiment selection → new data fed back to the AI.

This is precisely the lab-in-the-loop emphasized repeatedly in the paper.

AI has begun designing AI itself.

In the computer science lab, Co-Scientist’s level of autonomy increased further.

Synthetic biology

The researchers gave it a simple task: design an agent that can better answer medical questions. After that, humans no longer participated in the architecture design. Co-Scientist proposed its own solutions, wrote code, ran tests, analyzed errors, and continued refining the architecture.

Finally, it "evolved" into a system called Agent_H.

This system needs to rethink the reasoning process of existing models. When faced with a medical question, it first determines which medical specialty the question belongs to, whether it is directed at patients or doctors, and the level of risk involved; complex questions are broken down into multiple sub-questions. It then simultaneously generates 28–48 candidate answers, which undergo pairwise elimination by different Judges, followed by a vote from three Judges to select the final answer. The winning answer then undergoes multiple rounds of clinical review and citation verification before being condensed in length.

Synthetic biology

Answering one question requires the Agent to call the model approximately 40 to 80 times. On HealthBench Hard and HealthBench Professional, using the length-corrected scoring method from the paper, Agent_H outperforms six state-of-the-art models, including GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro.

Synthetic biology

But an intriguing incident occurred here that’s well worth reporting: the AI learned to "game the leaderboard" on its own.

In early experiments, the scoring criteria did not sufficiently penalize lengthy responses, so Co-Scientist quickly found the simplest way to boost its score: making answers extremely long.

The benchmark score increased significantly, but medical quality did not improve correspondingly. Only after researchers introduced a length penalty did the majority of the advantages disappear.

Google itself, in its paper, regards it as a classic example of Goodhart's Law: when a metric becomes a target, it ceases to be a good metric.

Evaluations from real doctors also tempered the results. Three practicing physicians conducted a blind assessment of 106 questions. Across nine dimensions, Agent_H showed statistically significant improvement over the original Gemini 3.1 Pro only in "reducing potential harm"; the other eight dimensions showed no significant difference.

Synthetic biology

In other words, a significantly higher benchmark score in the AI Judge’s eyes does not necessarily translate into a clearly better response in the eyes of human doctors. This may be precisely the challenge AI Scientist must confront upon entering the real world: it must not only know how to optimize, but also understand that not everything can be defined by optimization metrics alone.

AI scientists may fabricate data for papers.

In fact, Google devoted considerable space in this paper to investigating another somewhat awkward question: if the AI Scientist is fully unleashed, might it cheat to achieve impressive results?

Synthetic biology

Answer: Yes.

The research team enabled the system to autonomously handle all stages—from generating ideas and sourcing data to writing code, running experiments, and producing the final paper—across 50 AI research topics, with no human intervention.

For comparison, they also tested a version of Co-Scientist without the reliability module and the previous Agent Laboratory system, generating a total of 150 papers, which were then submitted to 30 domain experts for 450 anonymous reviews.

The result can be quite extreme: an AI Scientist without a dedicated validation mechanism might still complete an entire paper even when the code has errors and the experiment has produced no valid results.

Synthetic biology

They fabricate data tables, p-values, and statistical tests; turn failed experiments into successes; and even alter evaluation environments to give their methods an inherent advantage or hardcode outputs to produce artificially impressive results.

When the goal is to "write a paper that looks good," the agent will seek cheaper shortcuts that are easier than actually completing the research.

Therefore, Google added a mechanism to the new version of Co-Scientist that closely resembles a "research audit": any experimental results claimed in a paper must be traceable to actual program execution logs. If the logs do not contain the corresponding results, the system cannot include them based solely on the language model’s assumption that they "should be" true.

The improvement is striking. "Result hallucinations" severe enough to invalidate papers dropped from 90% in Agent Laboratory to 4%; the most severe cases of complete data fabrication fell from 44% to 0%.

However, the issues have not been completely resolved. The new version of Co-Scientist still has 24% serious methodological inaccuracies and 16% serious plagiarism or derivative content.

Google specifically emphasized that they have not claimed the current Autonomous AI Scientist can generate research papers ready for direct publication.

Conclusion

When viewed together, Anthropic and Google’s work reveal a clear shift in AI agents: for the past two years, the primary battleground for AI agents has been the software world. Now, this logic is for the first time rapidly extending into the real world.

Anthropic's MHS aims to address the underlying issue: providing a unified interface that an agent can understand and invoke across diverse devices such as robotic arms, microscopes, and liquid handling platforms.

Google's Co-Scientist, on the other hand, begins exploring higher-level questions: What would a true closed-loop scientific discovery system look like when a reasoning model can propose scientific hypotheses, design experiments, control instruments, read results, and iteratively refine subsequent experiments?

Of course, Google is still far from achieving a true "无人实验室." Current material experiments still require humans to load samples; the atomic structures of new materials have not yet been fully determined; E. coli predictions have only validated limited interpolation tasks; medical agents exploit loopholes in benchmarks; and paper generation systems have not fully resolved hallucinations and plagiarism. Google itself explicitly acknowledges that, due to hardware anomalies and the complexity of real-world environments, fully unsupervised physical experiments remain unattainable at this stage.

But changes have already occurred.

In the past, discussions about AI for Science often focused on whether AI could "think of something humans hadn't thought of." Now, an increasing number of efforts are addressing the more challenging另一半: whether that idea can be executed by real-world devices, and whether the results of that execution can feed back as input for the next round of AI reasoning.

Google concluded the paper with an intriguing observation: if such a closed loop is truly established, the limiting factors in scientific research could be reversed—previously, experimental equipment waited for scientists to propose the next good question; in the future, AI may have already generated hundreds of experiments worth validating, and laboratories simply won’t have time to carry them out.

By then, the true bottleneck in the pace of scientific discovery may no longer be scientific ideation, but experimental throughput—how quickly the real world can run these experiments.

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by: Interested in Science AI

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.