China's embodied intelligence data sector grows, with 25 startups securing over $17 billion in 2026

icon MarsBit
Share
AI summary iconSummary
China's industry trends in embodied intelligence data show strong momentum, with 25 startups securing over 17 billion yuan in funding by mid-2026. Companies such as LiberAI and Guanglun Intelligence are leading the way, providing robot-training data amid growing industry demand. The sector is divided into five approaches: real-robot data, simulation, UMI-based data, video distillation, and infrastructure platforms. Inflation data and investment flows remain closely linked, as downstream financing drives much of the data demand, creating a symbiotic yet fragile market.

At the end of July, LiberAI completed its Pre-A+ round of financing, raising hundreds of millions of yuan, with investors including 360 Group, China Development Bank Capital, Innovation Works, CMC Capital, Binfu Capital, and Sequoia China.

The company was actually founded on December 8, 2025, and is named "Jiangxian Technology," derived from the concept that "being able to relax is the primary productive force"; it has a registered capital of 1.46 million yuan and a team of fewer than 30 people.

IT桔子 found that this is already its fifth round of funding since its establishment; over the past six months, it has been raising funds rapidly, sequentially completing four rounds—from seed, angel, angel+, to Pre-A. According to reports, its latest valuation has reached RMB 5 billion.

Founder Liu Songming is a post-2000s graduate who earned Tsinghua University’s Department of Computer Science’s Top Prize Scholarship and ranked first in his major. He studied under Professor Zhu Jun and founded his company at age 23, immediately after completing his Ph.D. Previously, he led the development of RDT-1B, the first large-scale diffusion Transformer foundation model designed specifically for dual-arm robotic manipulation, with 1.2 billion parameters.

LiberAI targets human UMI data combined with world models—in simple terms, it sells the data fed to robots and the brain that processes that data.

According to IT Juzi, in the first half of 2026, 25 Chinese embodied intelligence data startups collectively raised over RMB 17 billion, indicating that "data" has become the most certain business in the embodied intelligence sector. (Excludes companies that also develop both embodied data models and robotic hardware, such as Xinghai Tu and Qianxun Intelligence.)

World Model

To view and download data on all companies and their funding in the embodied AI sector, visit the IT桔子 album page: https://www.itjuzi.com/album/766 (currently featuring 42 companies).

Among them, Guanglun Intelligence not only secured significant external funding but also secured orders worth RMB 550 million in Q1.

Before the gold miners strike gold, the people selling shovels make the first profits.

I. Unresolved Challenge: The Data Desert in the Robotics Industry

Why have large language models exploded in popularity? Because the internet offers an endless supply of text data.

But training the robot relies not on textual language, but on three-dimensional physical action data—such as picking up, putting down, walking, grasping, twisting, and tossing—that is nearly nonexistent on the internet.

Across the industry, the lack of embodied data is a major bottleneck.

Grand View Research predicts that the global data collection and annotation market will reach $17.1 billion by 2030.

The industry-recognized structure of embodied data is pyramid-shaped: at the apex are real-robot data, which are high-quality and directly deployable but also the most expensive; in the middle layer are simulated synthetic data combined with UMI data, offering moderate costs and unlimited scalability, yet suffering from the sim-to-real gap (friction and deformation approximation biases); at the base are internet videos and human behavior data, which are the most widely available but lack interaction details and are information-sparse.

Each layer has a cohort of companies relentlessly focused on it—this serves as a framework for understanding the current wave of embodied data entrepreneurship.

II. A Comprehensive Overview of Five Major Schools and Over Thirty Companies

School one: Physical remote control / Haptic school—treating data like a production line

This approach follows the most straightforward logic: build factories, set up robots, install motion capture systems, and mass-produce high-precision data on assembly lines.

The most aggressive is Pasini Perception.

The company leads the world in shipments of tactile sensors. In 2025, Pascin will complete the world’s largest embodied data collection factory, the Super EID Factory, in Tianjin, with an annual production capacity of 200 million data points. After raising over RMB 1 billion in its Series B round and achieving a valuation exceeding RMB 10 billion, the company has announced plans to build four additional super data factories in Suqian, Wuhan, Zigong, and Ganzhou.

Motion capture veteran is another branch of this line.

Dai Ruoli, founder of Noitom, previously led Noitom Technology to capture 70% of the global professional motion capture market. He later founded Noitom Robotics, which completed a several-hundred-million-yuan Pre-A++ funding round in June with a valuation of approximately RMB 3 billion. Beijing and Shanghai AI industry funds, Shenzhen Capital Group, CICC, and Kunlun Capital jointly invested, and the company unveiled its ModalityNet multimodal data platform.

QingTong Vision has adopted an open-source approach to build an ecosystem. In May, it released the world's largest open-source dataset of high-precision optical motion capture data, totaling 1,000 hours of human movement, with an annual production capacity of 500,000 hours. The motion capture data for Yuyu Technology’s Spring Festival Gala robot martial arts debut, “WuBOT,” was collected using its Project Decode system.

New entrants such as Lingsheng Technology are also positioning themselves.

Style Two: Simulation and Synthesis — "Printing" Data in the Virtual World

If real data is too expensive, can we create a virtually physically realistic world on a computer and batch-generate training data?

This route produced the world's first embodied data unicorn: Guanglun Intelligence.

In March, it completed a $1.5 billion A++ round; in May, it secured additional funding led by Ant Group; and in June, it raised another $1.5 billion, bringing its post-money valuation to over $2 billion (approximately RMB 15 billion).

Guanglun Intelligence's revenue increased tenfold in 2025, and new orders in the first quarter of 2026 reached RMB 550 million—exceeding last year's full-year total.

CrossDimension Intelligence is a representative of "generative simulation": its proprietary DexVerse simulation engine automatically generates massive volumes of scenario data in virtual environments; in July of this year, it announced a RMB 1 billion funding round, achieving a valuation exceeding RMB 10 billion and initiating its IPO, with RMB 100 million in revenue for the first half of the year and over 1,500 models successfully delivered for commercial use.

Among publicly listed companies, Qunxue Technologies, known as the "world's first space intelligence company," which debuted on the Hong Kong Stock Exchange in April this year, is leveraging its vast repository of 3D cloud design data to advance embodied intelligence.

The simulation platform Songying Technology raised hundreds of millions of yuan in combined Pre-A and Pre-A+ funding rounds in February and July, and collaborated with national and local entities to co-establish an innovative humanoid robotics center and develop synthetic datasets.

Style Three: UMI Ontology-Free / Portable Collection — Bypass the Ontology, "Replicate" the Person

There's a flaw in using remote-controlled robots to collect data—the movements captured aren't the person's true abilities, but rather compromised motions made to keep the robot in sync. So this approach bypasses the robot entirely, having people wear gloves, grippers, and wearable devices to perform tasks directly.

This is the most capital-intensive route this year.

On June 1, Jianzhi Robotics announced a series of consecutive funding rounds totaling hundreds of millions of yuan, led by Ant Group, Didi, and Delian Capital, with follow-on investments from Shunwei and BV (Baidu Ventures)—the largest funding round to date in the bodyless data domain.

Founder Chen Jianxing is the former Senior Director of Algorithms at Momenta. In its first year, the company has deployed solutions in thousands of real-world scenarios and secured a strategic partnership with Ant Group's Lingbo.

It Shihang proposed the Human-centric data collection paradigm and launched the SenseHub wearable collection solution; in April, it completed a $455 million Pre-A funding round, setting the record for the largest single investment in China's embodied AI sector, led by Hillhouse Capital, Sequoia Capital, and Meituan, and has initiated the "Embodied Data Spark Plan" with a goal of 100 million hours of data.

Tsinghua-affiliated Luming Robotics' FastUMI system has increased data collection efficiency by three times and reduced costs to one-fifth, earning continuous lead investment from Mitsubishi Electric and securing orders from top-tier clients.

In February this year, Zhìyuán Robotics spun off its data business into an independent company, Mifeng Technology, which raised hundreds of millions of yuan in seed and angel funding within just ten days, led by Sequoia Capital China, with a plan to achieve data production capacity in the millions of hours by 2026.

A new generation is also emerging: Xingyi Technology, spun off from Tsinghua University’s Department of Computer Science, is pursuing a path similar to NVIDIA’s EgoScale, developing first-person wearable capture systems with high freedom and millimeter-level precision; it has completed its seed round and secured a Pre-A round in July. Meanwhile, Professor Lu Zongqing from Peking University’s Zhi Zai Wu Jie is directly using internet video to pre-train general-purpose action models.

RiXi's spin-off, Qiongche Intelligence, has secured nearly a hundred orders for its "production-coupled" data collection system, CoMiner, and raised hundreds of millions of yuan in funding in June, with its self-developed world model set to be launched soon.

Another newly established company, Yuanche Taichu (OriginFlow), employs the NeuroScale paradigm to capture electrical signals from muscle contractions using a proprietary electromyography acquisition kit, then encodes and reconstructs hand posture, force, and haptic feedback through the PULSE foundational model. By directly interfacing with the native “intention-muscle-action” transmission pathway, it overcomes the limitations of UMI, such as visual occlusion and the absence of force and haptic feedback.

In its first year, the company has raised over RMB 500 million across its angel, strategic, and Pre-A1 funding rounds. The founder, Qin Shentao, is a 00s-born PhD graduate of Tsinghua University and a bachelor’s graduate of Harbin Institute of Technology.

School Four: Video Distillation / World Model School — Teaching AI to “Understand” the Physical World

This approach addresses the "turning waste into treasure" problem at the bottom of the data pyramid: there is an endless supply of human operation videos on the internet—cooking, handicrafts, repairs—but these videos lack the force, touch, and 3D trajectory information that robots need, rendering them as "visible but unusable" dead data.

The video distillation approach uses algorithms to infer motion trajectories and physical interaction information from 2D videos, transforming static data into trainable embodied data at a cost reduced to a fraction of a percent compared to real-world data collection.

The world model approach is more direct: first, let the AI learn the laws governing the physical world, then generate an infinite amount of training data "internally."

Shutu Technology's SynaData pipeline can batch extract multimodal embodied data, such as hand trajectories and object motion paths, from internet videos at an extremely low data collection cost. It has secured tens of millions in funding led by Oriental Fortune Capital, and its data is utilized by leading open-source models including Tsinghua RDT and AgiUniVLA.

Jiexia Shijie generated training data using the world model GigaWorld, improving the VLA model's performance by nearly 300% across three generalization dimensions. In June, it completed a B+ round of funding totaling RMB 1 billion, bringing total funding raised within three months to RMB 3.5 billion.

DeepWisdom has also invested in first-person human video data, raising over hundreds of millions of yuan across three funding rounds in the first half of this year. In May, DeepWisdom’s Z-WM World Model topped the global rankings in the WorldArena evaluation.

The LiberAI mentioned at the beginning of the article is also a hybrid of this approach and the world model赛道.

School Five: Data Infrastructure/Platform School — A data platform, also a "refinery"

This school of thought holds that the bottleneck in embodied data lies not only in “how to collect” but also in “how to use.”

Each robot has a different body, different sensors, and a variety of data formats; the raw data collected is like crude oil with mixed components—feeding it directly into models only makes them perform worse. Who will clean, align, label, and evaluate the data to refine the crude into standardized gasoline?

Moreover, training facilities, data collection equipment, and motion capture systems are capital-intensive—small and medium-sized teams simply can’t afford them. Therefore, there’s a need to build shared training facilities, establish data standards, and offer toolchains and platform services. Instead of betting on which specific data technology path will prevail, it bets that whichever path wins, it will have to pass through the infrastructure you’ve built.

Wuwen Zhike leverages the Deqing data collection and training facility, which features a closed-loop virtual-physical integration system, generating over a thousand hours of data daily. In April, it secured over RMB 100 million in funding and signed contracts worth hundreds of millions of yuan in the first quarter with clients including ByteDance and Wujie Power.

Yiren Technology completed two consecutive rounds of hundred-million-yuan funding in April and announced that it will achieve over one billion yuan in revenue and profitability in 2025, potentially becoming the first company in the industry to announce profitability.

Also invested in this “upstream” company, Zhiyu Jicheng, which specializes in data cleaning, alignment, and governance, are Lingchu Intelligence, Qiongche, Zhifangping, and other purely ontology-based companies established in December last year.

Publicly listed companies are also entering the scene, with leading data services provider HaiTian RuiSheng partnering with the Beijing Shijingshan Humanoid Robot Data Training Center to co-build an "Embodied Intelligence Data Training Ground."

Major companies are also entering the space: JD.com has launched an end-to-end infrastructure for embodied data, planning to mobilize 600,000 delivery personnel and riders for crowdsourced data collection, with a goal of accumulating 10 million hours of real-world video within two years; Baidu, meanwhile, has launched a “data supermarket.”

Three: Take a breath—is the demand side for data rigid?

2026 marks the year of large-scale adoption of embodied data; according to statistics, as of the end of April 2026, at least 90 data collection centers were in a "in use or under development/planning" state; of these, 64 are already operational, with the remainder under construction or in planning.

But the hotter the trend, the more you need to cool things down.

The embodied data business has an unspoken structural issue: both its demand and supply sides are funded by the same pool of capital.

Looking at the buyer list makes it clear—the large model teams, startup robotics companies, and ontology vendors in transition—most of these data purchasers haven’t yet turned a profit, and their procurement budgets come from recently received funding.

In other words, the revenue of data companies is essentially a reallocation of downstream funding: embodied AI robot companies raise capital, then use it to purchase data and tell compelling stories, which enable them to secure the next round of funding.

If the robotics company’s implementation falls short of expectations, data orders and funding will simultaneously recede. This is a symbiotic relationship—where success or failure is mutually dependent—rooted in downstream market demand.

Only when the robot company can truly achieve sustained profitability will this business become more solid and viable.

An industry-wide consensus holds that, in the long term, only two types of embodied data companies are likely to survive: one that becomes an industry-standard platform, controlling ecosystem-level tools such as simulation, data processing, and evaluation; and another that possesses cross-vendor data fusion and purification capabilities, continuously delivering high-quality, long-tail data after robots enter real-world environments.

This article is from the WeChat public account "IT Juzi" (ID: itjuzi521), authored by Wu Meimei.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.