Physical AI has taken over.Author and source: Dongjian Xin Yan She
Since the beginning of 2026, a popular term in the AI community has emerged—“Physical AI.”
At the beginning of the year at CES, Jensen Huang repeatedly stated, “The next wave of AI will be AI that operates in the physical world,” and Sun Yuchen recently made a high-profile declaration: “The era of virtual AI红利 has ended; physical AI is the biggest opportunity over the next three years.”
In the industrial sector, the prominent company Figure AI ignited online buzz with a non-stop five-day live stream of robotic sorting, while China’s AgiRobot announced the rollout of its 10,000th general-purpose embodied robot...
The statements from industry leaders and the real-world advancements in embodied intelligence have drawn the industry’s attention to this grand narrative of transitioning from virtual intelligence to physical execution. Yet, many still wonder: Is this so-called "physical AI" an inevitable turning point in technological development, or merely a well-packaged rebranding?
In 2026, the AI community is witnessing a surge in "Physical AI," with Jensen Huang stating that the next wave of AI will be AI operating in the physical world. Figure AI demonstrated its technology surpassing the lab-demo threshold through a five-day live stream of robotic sorting, while Agi Robotics achieved the production of its 10,000th general-purpose embodied robot. The core of this technology lies in enabling AI to achieve a closed-loop capability of "perceive-reason-act-feedback" in the real world. Key enablers include large language models granting robots comprehension, world models solving physical-world action challenges, and VLA models bridging the final mile from understanding to execution. Physical AI is transitioning from technical validation to commercial deployment, with over 110 billion yuan in funding raised since 2026, as competition enters the stage of mass production and delivery.Source: Dongjian Xin Yan She
01 From "Being Able to Chat" to "Being Able to Get Things Done"
Before answering the above question, let’s break down this slightly awkward technical term.
Physical AI, literally understood, is an artificial intelligence technology that deeply integrates AI with the physical world. However, at its core, virtual AI handles "thinking and communication," while physical AI must "perceive and act"—transforming it from an intelligent agent confined to a screen into one that can perceive, understand, and execute complex tasks in the real physical world.
Physical AI is a technology that "enables autonomous machines, such as robots and autonomous vehicles, to perceive, understand, and execute complex operations in the real physical world." Wang Xiang, an executive committee member of the China Computer Federation, systematically elaborated on this concept at the third China International Supply Chain Expo: "Physical AI means that AI systems possess a closed-loop capability of 'perceive—reason—act—feedback' in the real world."
In simple terms, previous AI could "chat," but today's physical AI can "do things." When AI steps out of the ChatGPT chatbox and into real-world environments like factories, warehouses, and homes, that's exactly what physical AI is designed to address.
This difference is particularly evident in the developments of two star robotics companies this year.
One is Figure AI from the United States, which used five consecutive days of live streaming to demonstrate that "robots can really work." The live stream, which began on May 14, showed three Figure 03 humanoid robots taking turns on the production line to sort packages. The robots’ tasks included scanning barcodes, picking up packages, reorienting them, and placing them with the barcode facing down onto the conveyor belt.
During the livestream, a robot worked continuously for over 33 hours, handling more than 40,000 packages. Founder Brett Adcock stated that the robot operated in "fully autonomous mode" using the company's latest Helix 02 model.
The significance of Figure AI's live stream lies not only in demonstrating its technological capabilities but also in using real-time footage to tell the world that physical AI technology has crossed the threshold beyond "lab demonstrations." A company broadcasting live for days, showing robots operating continuously on a production line without major issues, is itself a powerful technological declaration.
Zhìyuán Robotics in China also conducted a similar live stream, placing its Zhìyuán Elf G2 robot on the MMIT (Multimedia Integration) production line at the Nanchang Longqi Technology Park to work alongside humans. Real-time test data from the livestream showed that the robot operated continuously for eight hours with zero major anomalies and achieved an overall operational success rate of over 99.5%. Each individual task took only 18–20 seconds, enabling the robot to complete 310 units per hour—equivalent to the workload of two sequential processes handled by a single robot.
Building on its progress with Figure AI, Agi Robotics announced in March that it had delivered its first 10,000 units of the world’s first general-purpose embodied AI robot, achieving the leap from 5,000 to 10,000 units in just over three months—from December 2025 to March 2026.
Beyond delivery volumes, AgiBot revealed that the company aims to achieve RMB 10 billion in revenue by 2027. Based on past development patterns in cutting-edge industries such as new energy, autonomous driving, or semiconductors, a company less than two years old that has achieved mass production and delivery at the ten-thousand-unit scale and set a billion-yuan revenue target is nothing short of phenomenal in the hardtech sector.
The two companies above have demonstrated with concrete data and real-world scenarios that physical AI no longer relies on remote control or pre-programmed scripts to "perform," but instead can autonomously complete complex tasks in real environments.
More importantly, Zhiyuan has become the first to surpass the threshold of 10,000 units delivered, linking its mass production capacity with existing orders, signaling a turning point in this sector from “technology validation” to “commercial realization.” In other words, the feasibility of physical AI is no longer in question—the real competition has now entered the deeper waters of “usability” and “economics.”
02 Technological Drivers Behind the Surge of Physical AI
So, the question arises: why did physical AI suddenly explode this year? Looking back, beyond genuine commercial demand, a series of technological breakthroughs have been the biggest driving force.
First, large language models (LLMs) have endowed robots with "understanding capabilities." Traditional robots rely on deterministic code and rule-based programming, essentially requiring engineers to pre-write a "script" that dictates every action the robot must strictly follow. This approach has a major flaw: even minor changes in the robot’s operating environment necessitate rewriting the code, resulting in poor robustness and making it difficult to cross the threshold into commercialization.
However, after Google attempted to integrate LLMs with robotic physical execution and successively launched embodied multimodal large models such as Google PaLM-E and RT-2 in August 2023, robots gained the ability to automatically decompose complex tasks into multiple steps and execute them via natural language instructions, marking the LLMs' transition from “conversational understanding” to “physical execution.”
At CES 2026, Jensen Huang highlighted the essence of this technological evolution: Physical AI is fundamentally a transfer of underlying control. When physical AI crosses the tipping point of technological advancement, control shifts from deterministic code written by humans to neural networks with generalization capabilities that understand physical laws.
At this point, the robot is no longer just “executing code,” but has gained the ability to “understand instructions and autonomously plan its actions.”
If large language models solved the problem of "understanding," then world models solve the problem of "acting in the physical world." The core of world models is enabling AI to develop an internal understanding of the laws governing how the physical world operates.
NVIDIA's Cosmos, a foundational model platform for physical AI unveiled at last year's CES, became a landmark event; its core capability is generating physics-compliant motion data from text or images, enabling developers to accelerate the development of physical AI for autonomous vehicles, robots, and video analysis AI agents.
According to NVIDIA, Cosmos is trained on over 20 million hours of real-world data, significantly reducing the complexity of simulation and model training. With this world model, AI systems can conduct massive simulations in virtual environments before transferring their learned behaviors to the real physical world.
The ultimate capability of a robot is not "seeing" or "understanding," but "doing correctly." The emergence of Vision-Language-Action models enables robots to simultaneously process visual input, language comprehension, and motion control, achieving a closed loop of "seeing is doing."
In September last year, DeepMind released its next-generation multimodal embodied intelligence large model, Gemini Robotics 1.5, claiming it to be the world’s first thought-oriented model optimized for embodied reasoning; NVIDIA, meanwhile, launched the open-source Isaac GR00T N1.6 model designed specifically for humanoid robots, enabling full-body control.
Meanwhile, the Beijing Humanoid Robot Innovation Center has open-sourced the embodied cerebellum large model XR-1, which has become China’s first model compliant with the national standard for embodied intelligence. Trained on over one million data points, it can perform complex dual-arm manipulation tasks such as picking up, placing, pushing, pulling, and rotating.
At this point, physical AI has gathered all the foundational technical capabilities required for real-world deployment: LLMs enable machines to understand human intentions, world models allow machines to predict physical consequences, and VLAs bridge the final gap from “understanding” to “doing” correctly. Together, these three components equip robots with the foundational ability to autonomously execute tasks in open environments for the first time.
Of course, dexterous manipulation still faces bottlenecks—fine control of arms and hands continues to present many unresolved challenges. In other words, physical AI has received its ticket to enter factories and perform labor, but to truly step into homes and serve tea or perform delicate tasks, it must overcome the qualitative leap from coarse movements to refined operations.
03 From Technical Vision to Delivery Capability
It is important to understand the past and present of physical AI; now, the embodied intelligence industry must address the question of which core dimensions future competition will revolve around.
Drawing lessons from the development of autonomous driving, the battle for data cannot be avoided—and embodied intelligence, which follows a similar logic, cannot avoid it either. Generally, whoever possesses higher-quality training data holds the power of influence.
In the industry today, NVIDIA has taken the lead in establishing a moat for world models using Cosmos; its model, trained on over 20 million hours of real-world data, is difficult to replicate quickly. Meanwhile, Zhìyuán has achieved mass deployment of 10,000 robots, giving it the capability to collect real, feedback-driven data—a capability widely regarded in the industry as a data moat.
It should be noted that the data required for physical AI competition is not simply about who has the largest volume, but rather about the synergy between synthetic and real data.
Relying solely on real data presents challenges of scale and hardware wear costs, while over-relying on synthetic data creates a sim-to-real transfer gap. Beijing Humanoid Robotics Innovation Center’s “cross-data-source learning” solution emerges from this insight, enabling robots to train using vast amounts of human video data, significantly reducing training costs while improving training efficiency.
This makes it much clearer: whoever can truly establish the complete closed loop of “synthetic data training—real data fine-tuning—real-world feedback” will gain a competitive advantage in this race.
After resolving the data issues, efficiently integrating physical AI with virtual AI became the key to advancing physical AI further.
When discussing physical AI, one direction often overlooked is that physical AI and virtual AI are not opposing concepts. From a technical architecture perspective, a complete physical AI system can generally be divided into three layers: the bottom layer is the perception layer (sensors, visual recognition), the middle layer is the cognitive decision-making layer (AI reasoning), and the top layer is the action execution layer (mechanical control).
Virtual AI primarily handles the middle layer, while physical AI must connect the entire chain from perception to execution.
NVIDIA’s full-stack solution of “chip + model + tools” embodies this approach: the Jetson Thor edge computing platform provides computing power, the GR00T model delivers intelligence, and the Isaac platform offers a development toolchain. Looking ahead, whoever can achieve deep integration of software and hardware will not only close the loop of physical AI—from “brain” to “limbs”—but also build a strong technological moat.
Lastly, regarding the commercialization of physical AI, three years ago, capital's enthusiasm for the robotics sector was driven by "technological vision"; now, the investment community has adopted a more pragmatic evaluation standard: delivery capability.
Media statistics show that in 2025, China’s embodied intelligence sector raised a total of RMB 73.5 billion across 744 investment and financing events. Since 2026, an additional RMB 37 billion has been added, bringing the cumulative total beyond RMB 110 billion. Yet beneath this flourishing surface, capital flows have undergone a visible structural shift.
In May 2026, Tianji Intelligence completed a B-round financing of RMB 1 billion, with its key milestone being over 10,000 units in hand orders for Q1, serving 45 robotics companies.
Zhongke Fifth Era has simultaneously secured a hundreds-of-millions-yuan Series A funding round and disclosed that it has secured overseas orders worth hundreds of millions of yuan.
During the funding rounds of Weitai Power and Lu Ming Robotics, industrial investors such as SAIC Shangqi Capital and Mitsubishi Electric have successively joined, aiming to link production capacity with robot delivery capabilities.
In contrast, the American humanoid robotics startup Cartwheel Robotics, despite having a technological vision, lacked orders to sustain it and declared bankruptcy in March 2026.
Positive and negative examples show that capital no longer pays for flashy demos—it only pays for proven, scalable delivery capabilities.
04 Conclusion
The sudden surge in popularity of physical AI is actually a natural progression.
Of course, some industry insiders view "Physical AI" as merely a new concept packaged by capital markets, essentially a natural evolution of embodied intelligence and robotics technology. However, there is no denying that the rise of Physical AI clearly signals that the AI industry is transitioning from "virtual intelligence" to "physical execution"—a historical trend that is irreversible.
In the latest round of competition, Figure AI showcased its capabilities through live streams, Agi Robotics built industrial barriers through mass production and delivery, and NVIDIA constructed a platform ecosystem with Cosmos and GR00T... The next question is: which company will become the OpenAI of physical AI? And which application scenario will be the first to experience its “ChatGPT moment”?
