DeepSeek The official release of V4-Flash has just shaken the global general-purpose model community, and China’s life sciences version of DeepSeek has quickly followed suit.
Recently, Jindu Life Science—founded by four Oxford University returnee scholars—successfully published its self-developed life sciences-specific GeneLLM multi-omics large model in the internationally renowned academic journals Nature Communications and Advanced Science.
As the world’s first multi-omics large model pre-trained directly on raw omics data, GeneLLM is a groundbreaking model following Google’s AlphaFold and Stanford University’s EVO 2, filling the gap in China’s foundational large models for life sciences and earning its reputation as China’s answer to DeepSeek—enabling AI to begin understanding multiple “languages” of life and the entire biological “system.”



Predict the next piece of life information, just like predicting the next token.
Disease identification is just one application scenario of GeneLLM; what JinDu Biosciences truly aims to build is the 'Claude Code' of the life sciences field.
ChatGPT, Claude, DeepSeek At the core of large language models is next-token prediction.
GeneLLM follows a similar idea, but it predicts biological information rather than text.
The four bases in RNA sequences—adenine (A), uracil (U), guanine (G), and cytosine (C)—serve as the fundamental tokens through which GeneLLM understands the language of life.
Traditional bioinformatics analysis typically relies on gene annotation, sequence alignment, and manually defined labels. While this approach is accurate, it inherently limits the model’s scope of understanding and may overlook numerous unknown biological signals hidden within the raw data.
GeneLLM takes a different approach by learning the laws of life directly from raw, unprocessed sequencing data.
It uses multi-omics raw data, such as transcriptome, proteome, and metabolome, as training data to enable the model to autonomously discover disease-related patterns.
GeneLLM has completed pre-training with a 1.5 billion parameter model on 3.5 trillion base sequences, and the XLarge version has achieved pre-training with a 30-billion-parameter model, continuously expanding its technological edge.
As the world's first multi-omics large model pre-trained on raw sequencing data, GeneLLM consists of two stages:
(1) Unsupervised Pretraining and Prototype Mining
(2) Patient-level disease fine-tuning

GeneLLM first needs to address an issue:
How can complex biological data be converted into language that AI can understand?
In natural language processing, the BPE algorithm splits sentences into tokens; GeneLLM, however, divides RNA sequencing fragments of approximately 150 bp into life tokens using a sliding window of seven bases (7-mer).
Subsequently, the model uses the Transformer architecture to directly predict the next base without gene annotations or manual labels.
This means AI does not first learn from a human-curated "dictionary," but instead learns directly from raw signals of life.
During training, GeneLLM processed approximately tens of trillions of RNA reads on a cluster of hundreds of NVIDIA A100 GPUs.
At this point, GeneLLM's generalization and other capabilities begin to emerge.
As a major innovation breakthrough in the life sciences, GeneLLM, a domestically developed life sciences model, can be applied to multiple fields including drug discovery, precision medicine, synthetic biology, environmental monitoring, microbiology and bio-agriculture, and protein and molecular design—it is one of the few large models worldwide to have achieved real-world deployment.

Only after examining GeneLLM’s foundational innovations in “data, architecture, and training” can we truly understand why it represents China’s “biological version” DeepSeek ".
Silicon Valley has piled up tens of thousands of H100s to build models with trillions of parameters, DeepSeek By leveraging algorithmic innovation to enhance computational efficiency, GeneLLM similarly uses a small force to move a great weight, unlocking billions of life science data points.
If AlphaFold enabled AI to first understand the "structure" of life, and EVO2 allowed AI to begin comprehending life's "code," then GeneLLM seeks to further understand life's "system."

Even more impressive is the efficiency. Traditional methods rely on 6 Gb deep sequencing, which is costly and difficult to implement. GeneLLM maintains an AUC > 0.8 even at an extremely shallow depth of just 1 Gb, reducing costs by 83%.
This means making precision medicine accessible to all is now truly within reach.
A new scientific paradigm is emerging: enabling AI to learn from biological data, ushering life sciences into a new era of predictability, computability, and scalable exploration.
Not just a model, but also intelligent infrastructure for the last mile.
But Jin Du Shengke's strategy extends beyond foundational models.
Building on the GeneLLM multi-omics large model as the foundation for life science understanding, JinDu Bioscience has further developed an execution system that connects AI intelligence with the physical world. Through the Harness intelligent experimental execution layer and the DBTL (Design-Build-Test-Learn) data loop, it achieves a complete AI-for-Science闭环—encompassing the understanding of biological principles, generation of scientific hypotheses, automated experimental validation, and continuous iteration and optimization.
As Liam Fedus, former OpenAI VP and head of post-training, stated, current LLMs have exhausted the limited text and code available on the internet; significant scientific breakthroughs moving forward must rely on experimental iteration.

In other words, AI cannot discover new knowledge solely by reading content written by humans—it must conduct its own experiments.
But here's the question—
Internet services inherently have APIs, and online information is inherently digital, but scientific equipment comes from different manufacturers with varying protocols, making it both complex and expensive—no one can tear it down and rebuild it in a few months.
Today, most laboratories are designed for humans: instrument panels, pipetting motions, sample conditions, and on-site judgments have rarely been converted into machine-readable signals.
AI has entered the lab; it's no longer just about model capabilities, but also about building a new infrastructure.
Therefore, AI has entered the lab, and now it’s not just about model capabilities—it also requires a physical harness: turning the lab into a system that is compilable, schedulable, observable, and traceable.
This is the true last mile of AI4S.
BioFord Harness, enabling the lab to start "doing it themselves"
To this end, Jindu developed a system called BioFord Harness.
This is an infrastructure that connects AI with physical laboratories.
Instead of having robotic arms mimic human hands, they transformed the laboratory into a system that is programmable, schedulable, observable, and traceable.
During this process, the physical harness must accomplish at least three things:
1. Compile scientific intentions or experimental DSL into instructions executable by different devices;
2. Manage scheduling, resource allocation, and security constraints across multiple devices, and handle exceptions;
3. Feed back experimental results, device logs, and environmental parameters as inputs for the next round of model and experiment design.
Scientific question → AI understanding → Experimental design → Equipment scheduling → Execution → Data feedback → Model optimization → Next cycle.
A flywheel of scientific experimental data has thus been set in motion.

Five intelligent agents turn the research process into an assembly line.
The BioFord Agent Embodied Intelligence Research Platform connects downward to the physical layer of laboratories and upward to the cognitive layer of scientists, bridging the gap between reasoning and execution with physical AI.
At the cognitive level, the BioFord Agent consists of a collaborative network of five intelligent agents that streamline the entire life science research process, increasing research efficiency by several times.
Literature search agent
Experiment Design Agent
Scientific Agent
Experiment Scheduling Agent
Data Analysis Agent
For example, a literature search agent can quickly retrieve and read vast amounts of literature, complete a literature review, and assist in formulating hypotheses.

The experiment design agent enables research teams to reduce their experiment design cycle from months to just one week.
Jindu Life Science's experimental scheduling agent, built on the Universal Instrument Abstraction Layer, breaks down protocol barriers between diverse heterogeneous devices.
Whether it's a PCR instrument, microplate reader, flow cytometer, or automated liquid handling workstation, all can be centrally managed and coordinated. The system features an integrated dynamic scheduling algorithm that automatically batches scheduling, real-time conflict avoidance, and fully records experimental parameters to create an auditable, traceable chain.

This means that AI is no longer limited to the stage of "providing advice" but has truly entered the laboratory, operating equipment and executing tasks as a reliable "research assistant."
Failed data may be more valuable than successful data.
The deeper value of this system lies in turning every experiment—whether successful or not—into data that the system can process.
In a traditional lab, a failed experiment might simply be recorded as “results did not meet expectations,” and the expertise resides only in people’s minds—when they leave, the knowledge is lost.
In the BioFord system, every failure becomes valuable training data—logic behind parameter selection, environmental conditions recorded, proof of incorrect paths—all of it accumulates as part of the system’s “experience.”
Next time, the AI will know: This path leads nowhere.
Interestingly, in this space, the winner isn't the one with the most computing power or the largest model.
Computing power can be purchased, but research data cannot.
Kim Young-sung, Founder and CEO of JinDooSang, said: “During our R&D process, we gradually realized that AI for Bioscience is not simply about stacking models and data. For those conducting experiments, the ultimate challenge remains solving the dilemma of getting results that can’t be executed or that are incorrect even after execution.”
This is the most important and hardest-to-replicate moat in the AI4S sector.
Four Oxford graduates, including a senior fellow of Luo Fuli's same academic lineage.
In 2022, Kim Young-seong faced a choice the moment he earned his Ph.D. in biomedical engineering from the University of Oxford.
His advisor is Hagan Bayley, a Fellow of the Royal Society and founder of Oxford Nanopore, a UK-listed leader in third-generation sequencing, where the tradition of turning research into real-world applications runs deep. Staying in the UK was a clear and straightforward path.
But he chose a different path—he packed a “prototype technology” from the lab into his suitcase and brought it back home. Joining him were three Oxford alumni: Dr. Deng Siwei in biology, Associate Researcher Sha Lei from the Computer Science Department (a fellow student of Luo Fuli, head of Xiaomi’s large model team, and a Ph.D. graduate from Peking University’s Computer Science Department), and Zhou Tianyao, an expert in product commercialization. The four possessed complementary expertise in bioengineering, artificial intelligence, computational biology, and business operations—each indispensable. They had previously collaborated on joint research projects, combining AI with transcriptomics to enable disease prediction and detection. Like puzzle pieces, they fit together perfectly, enabling highly interdisciplinary research.
The company name “Jin Du Sheng Ke” draws “Jin” from Oxford, symbolizing their origin at elite academic institutions such as Oxford’s top laboratories, and “Du” signifies guiding others—a vision they believe is the ultimate goal of AI for Science.
Currently, Jin Du Life Science's BioFord Agent physical AI research platform has been implemented at several renowned universities in China, significantly reducing research cycles from months to just one week.
The wind is rising: Four funding rounds in one year—a real case of capital FOMO
However, walking a path that has never been taken before inevitably brings loneliness.
In the early stages of entrepreneurship, challenges were abundant. At that time, AI startups were thriving, but there were very few attempts to combine AI with biosciences. “Those who understand AI may not understand biosciences, and most who understand biosciences don’t know AI.”
Kim Young-sung admitted: “Investors once couldn’t understand what we were doing.”
The team chose the "most difficult path": entering through biological foundation models, a direction pursued by only a handful of companies worldwide.
In 2025, a turning point emerged.
At that time, the State Council issued the "Opinions on Deeply Implementing the 'AI+' Initiative," which included AI for Science, marking the beginning of a new trend in this sector.
Jindu Shengke achieved the feat of completing four funding rounds in one year. The company’s primary funding timeline is considered an industry benchmark.
Angel + Seed Round: Led by Sequoia Capital China Seed Fund;
Pre-A+ Round: Led by Chuangdongfang Investment in a seven-figure funding round;
Pre-A+ Round: Secured a seven-figure investment from Nanshan Strategic Emerging Industries Investment;
Series A: Led by Gao Tejia Investment with nearly RMB 100 million.
Teng Yuhang, Executive Partner at Gaotai Investment, said: "Jindu Bioscience has transformed foundational life sciences research into a subscription-based, scalable 'compute power + experimentation' infrastructure."
But Kim Young-sung’s vision is even broader: “We’re not satisfied with just selling software—we aim to build an intelligent operating system for the life sciences. Just as Intel defined computing power in the PC era, we want to define a new paradigm for R&D in life sciences in the AI era.”
From Oxford Labs to Shenzhen, from being misunderstood to receiving heavy investment from top-tier capital, the story of these four Oxford graduates is not just a创业 legend, but a profound questioning of “What can we do for humanity?”
When AI learns to "pull all-nighters" in research, scientific discovery may no longer rely on the accidental brilliance of geniuses, but become an expected and inevitable outcome. They have only just begun this journey.
Industry Map: Jin Du Pioneers a Lightweight Physical AI Approach
On the larger AI for BioScience map, Jin Du is not alone, but its entry point is distinctly different.
The first category consists of AI scientists in the digital domain.
Biomni, incubated at Stanford and commercialized as Phylo, offers more than 150 specialized tools that automate literature reviews, hypothesis generation, and bioinformatics analysis.
FutureHouse, backed by Eric Schmidt, is dedicated to building AI scientists capable of autonomously generating hypotheses and writing papers.
However, although powerful, they remain confined to the digital world.
The second type is the full-stack autonomous approach.
Jingtai Technology has deployed over 300 workstations globally using "AI + robotics."
Lila Sciences, incubated by Flagship Pioneering, has raised $550 million to enable AI to fully take over the design, execution, and redesign of experiments, with the goal of achieving a "scientific superintelligence."
But all of these are too expensive.
The third category is end-to-end managed routes.
Insilico Medicine is directly integrating AI into its proprietary drug development pipeline; its first AI-generated drug has entered Phase III clinical trials, but they are "builders of cars," not "road repairers."
Jindu Shengke’s choice is to focus on building physical harnesses as the infrastructure for the "last mile" between models and physical experimental systems.
Instead of competing in foundational layers, end-to-end pipelines, or being purely digital-world AI scientists, focusing solely on that final kilometer is undoubtedly a more lightweight approach.
Kim Young-seong put it clearly: “The true watershed for AI for Science is not how closely the model mimics a scientist, but whether the laboratory can begin to function like a continuously learning system.”
Moreover, in the global AI for Science field, Jindu is not merely "doing the last mile."
More precisely, it enters through the last mile to compete for control over the entire research workflow.
Compared to purely digital AI scientists, it can interact with the physical world; compared to building capital-intensive scientific facilities from scratch, it has the opportunity to take over customers’ existing laboratories; compared to end-to-end AI drug discovery companies, it does not have to bet its future on a single clinical pipeline.

Its true moat will not be parameter count, but the accumulated experimental trajectories, equipment interfaces, failure experiences, and cross-laboratory execution networks.
As Kim Young-sung said, the tens of thousands of signaling pathways and unknown reaction mechanisms within living organisms are fascinating. "The richness of biology is matched only by the beauty of AI for Science."
Using AI to intelligently explore the mysteries of life, this group of Oxford scholars has only just begun their journey among the stars.
This article is from the WeChat public account "New Intelligence Yuan," authored by New Intelligence Yuan; edited by Aeneas KingHZ.
