Experts Discuss Physical AI as the Next AI Paradigm at Tencent WAIC Event

iconMetaEra
Share
AI summary iconSummary
At the 2026 World Artificial Intelligence Conference (WAIC), hosted by Tencent News and Tencent Technology, experts discussed the emergence of physical AI as the next major AI paradigm. Panelists such as Huang Xiaohuang of Qunheu and Luo Yihang of Shenshu highlighted the transition from digital to physical AI, including world models and spatial intelligence. They underscored the necessity for AI to function beyond screens, with some predicting a "ChatGPT moment" for physical AI within three to five years. The event also attracted attention from the AI + crypto news sector, with discussions addressing real-world assets (RWA) and how physical AI could transform digital-physical integration.
AI is transitioning from digital agents to physical agents.

Article author and source: Tencent Technology

An increasing number of professionals are beginning to ask.

As the competition among large language models evolves into a super capital-intensive arms race, an increasing number of industry professionals are asking the same question: Where is the next paradigm for AI?

The answer is moving from "within the screen" to "beyond the screen."

On the evening of July 17, during the 2026 World Artificial Intelligence Conference (WAIC), the "Tencent WAIC Night" was held in Shanghai, hosted by Tencent News and Tencent Technology, with special support from Tencent East China Headquarters and Tencent Cloud Intelligence.

At the event, Chen Yu, Managing Partner of Yunqi Capital; Huang Xiaohuang, Co-founder and Chairman of Qunxie Technology; Luo Yihang, Co-founder and CEO of Shengshu Technology; and Zhang Zhizheng, Co-founder and Large Model Lead at Beijing Galaxy General Robotics, jointly participated in a roundtable discussion titled “What Comes After Agentic AI?” focusing on spatial intelligence, world models, and Physical AI.

From left to right: Chen Yu of Yunqi Capital, Huang Xiaohuang of Qunxiao Technology, Luo Yihang of Shengshu Technology, and Zhang Zhizheng of Galaxy General.

Huang Xiaohuang opened by presenting a perspective: he believes that pure software agents without industry expertise struggle to maintain an advantage, as large models can easily replicate them. In the future, tools, hardware, data, and large models must work together to continuously amplify the value of AI products. GroupCore’s focus is on “generating an interactive world with large models”—a direction that lies beyond the reach of large language models: Physical AI.

In January this year at CES, NVIDIA founder Jensen Huang repeatedly emphasized that AI will evolve from perception and generation to agency, ultimately reaching "physical AI" capable of understanding the physical world, and asserted that the "ChatGPT moment for physical AI" is imminent.

Following the thread of the "physical world," Luo Yihang summarizes it as the "universal model of the physical world": Just as the scaling laws of language models have validated generalization, the physical world will also soon reach its own "ChatGPT moment"—he predicts that embodied generalization could arrive within 3 to 5 years.

Zhang Zhizheng’s statement is the most concise—moving from Digital AI to Physical AI, the learning paradigm will shift from offline learning to online learning: “First learn to interact, then continuously learn from interaction.” Interaction is no longer the goal, but rather a data flywheel that drives capability growth.

The shift from agentic AI to physical AI is being repriced by capital and policy.

In early 2026, the World Economic Forum assessed that AI is “beginning to operate as a physical system within the real economy”; Deloitte’s “2026 Technology Trends” explicitly states that physical AI is ready for mainstream deployment. Data further supports this: the global market size for embodied intelligence is projected to reach approximately $4.44 billion in 2025 and is expected to grow to $23 billion by 2030, with China’s industry scale potentially surpassing one trillion yuan by 2035.

Although all parties share a cognitive consensus on the direction, their paths diverge significantly—the focus is precisely on the "world model."

Huang Xiaohuang believes the world is driven by invisible laws such as gravity and friction, and that vision is merely superficial; thus, he is more optimistic about returning to physical simulation. Currently, Qunxue Technology’s exploration of world models is gradually converging toward spatial intelligence. At the other end of spatial intelligence, Li Feifei offers a different vision through World Labs’ first commercial world model, Marble, whose valuation has surged from $1 billion to $5 billion in just over a year. Luo Yihang advocates for integration, arguing that any single path will eventually hit a bottleneck; Zhang Zhizheng breaks down the “world model” into three indispensable capabilities: modeling how actions alter states, using modeling to enhance strategy learning, and possessing general evaluative ability.

On the implementation timeline, each of the three has a different focus: Huang Xiaohuang views “digitizing physical-world information at low cost and at scale” as the primary challenge; Luo Yihang admits that current models and data are “still not there,” but has already seen precursors similar to the transition from GPT-2 to GPT-3.5; Zhang Zhizheng’s assessment is the most aggressive—he believes the logic of physical AI is “pursuing comprehensive generalization within limited scope,” and therefore, the arrival of physical AI agents may significantly shorten the cycle compared to digital AI.

The answer to the paradigm shift may not come from any pre-determined technological path, but will gradually emerge through this ongoing debate.

Below is the transcript of this panel, with adjustments and deletions made without altering the original meaning.

The Next Paradigm for AI

Chen Yu: Thank you to Tencent for the invitation. Over the past year, the development of AI can be described as progressing by leaps and bounds—whether in the dramatic advancements of AI capabilities themselves or the performance of major large model companies in capital markets. However, AI has long moved beyond language models, extending into embodied intelligence, video generation, and spatial intelligence.

Today, we are honored to have the CEOs and Chairmen of companies that are either about to go public or already listed, specializing in spatial intelligence, video generation, and embodied intelligence, to discuss: What is the next AI paradigm after Agentic AI? First, please introduce yourselves briefly from left to right, and answer in one sentence: What do you believe is the next AI paradigm?

Huang Xiaohuang: Hello everyone, I’m Huang Xiaohuang, co-founder and chairman of Qunxue Technologies. Our company is positioned in spatial intelligence. When we first started, I used GPU clusters for spatial rendering. Later, as times evolved, we began doing “inverse operations”—deriving structured data from rendering results or real-world outcomes—and thus accidentally entered this field.

Over the past few years, I’ve been reflecting on one thing: large language models have entered a hyper-capital-intensive race, so I often wonder which models lie outside the reach of large language models.

After observing the development of large models for over a year, my impression is that pure software-based agents without industry-specific expertise will face increasingly weak competitive moats. Since large models can essentially replicate software or agents directly, relying solely on tools makes it difficult to establish long-term competitive advantages. My preliminary assessment is that, in the future, tools, hardware, data, and large models will need to be deeply integrated.

Where do we go from here in this era?

After spending two to three years researching and exploring, we concluded that large models are essential—but they must go beyond the reach of large language models. So initially, we trained four or five different models, eventually converging on spatial intelligence: how to generate an interactive world that is difficult to precisely describe in language. This is somewhat similar to Li Fei-Fei’s Marble, but we felt that purely static scenes weren’t enough—we needed the ability to interact to some degree.

In summary, it’s about having a large model generate an “interactive world”—this is the direction we’ve chosen, as it lies beyond the reach of large language models.

Luo Yihang: Hello everyone, I’m Luo Yihang, Co-founder and CEO of Shengshu Technologies. Shengshu focuses on multimodal generation and world models. What will the next paradigm be? I believe it will be something that no one currently envisions, hasn’t converged on, and hasn’t reached consensus about. My conclusion is: a universal model of the physical world.

Why? First, the core paradigm of this generation of models is the Scaling Law, which has already been validated for its generality and adaptability in language models. I believe a similar paradigm will emerge in the physical world to address interaction and generalization challenges in physical environments. It must achieve generalization and reasoning: first, there will be architectural innovations—just as Transformer revolutionized language models, we are exploring architectures like Diffusion Transformer or MoE hybrids for world models; second, there will be innovations in data—relying solely on current embodied sensor data, I believe, is not the optimal path.

In the future, we can imagine that robots of various physical forms and architectures will be like Agents and WorkBuddies in our digital world—they can autonomously plan, collaborate, and complete long-term tasks, extending the concept of Agents into the physical world. While the physical forms of these robots will certainly vary, their intelligence will be relatively general-purpose.

We believe the world model will have its own "ChatGPT moment," giving rise to such generality and adaptability; when embodied world models reach a certain level of intelligence, combined with advancements in the body and cerebellum, I believe the era of embodied generalization could emerge within three to five years.

Zhang Zhizheng: Hello everyone, I’m Zhang Zhizheng from Galaxy General, co-founder and head of the large model team. When it comes to the next shift in AI, I believe it’s very clear: the transition from Digital AI to Physical AI—from building AI models for the digital world to building models that can interact with the physical world. This is precisely why embodied intelligence is receiving so much attention in both academia and industry today.

During this process, the learning paradigm undergoes a fundamental shift. In the past, when we built GPTs and large models for the digital world, we often relied on offline learning. However, to build a model that can reliably interact with the physical world, we must first transition from offline learning to online learning—learn to interact first, then continuously learn from those interactions.

Under this paradigm, interaction is not the ultimate goal of learning, but a means to turn our AI into a flywheel of capability and data, continuously propelling itself forward.

Driven by this data flywheel, I believe we will move from the digital agents widely used today to physical-world agents capable of handling various tasks in the physical world. This will bring even greater transformation to the entire industry—across both development and usage—as well as to the way we live and work.

02 "World Models": The Battle Between 3D Simulation and Pure Video

Chen Yu: Just now, several of you mentioned that physical AI may be the next paradigm of AI development, and all of you also brought up world models. I’ll therefore move the topic of world models forward a bit.

There is currently much confusion surrounding world models—I might receive five different business proposals for world models in a single day, yet everyone’s understanding and definition of a world model varies. How do people interpret world models? What exactly are they?

Huang Xiaohuang: I'll go first. Internally, we've also been continuously researching world models, and overall, there are essentially two approaches: one is the traditional 3D simulation route, using structured data for simulation; the other is the pure video training route.

The former believes that the workings of the world lie beyond visual perception—that the motion of objects arises from the interaction of forces such as gravity and friction, or from some form of "causality." Therefore, they hold that the world is composed of invisible factors, and that relying solely on visual observation is unreliable; vision, they argue, merely "deceives humans."

The latter argues that, as human perception of the world relies primarily on vision, a purely video-based model should work just as well, provided there is sufficient data and upgraded training algorithms.

Our past expertise has been in parametric models and structured data.

I can’t currently prove which direction is definitely right or wrong; we consider both possibilities viable, but we personally favor the first—we believe the world is composed of countless physical laws and constraints, and ultimately, how the world operates must be understood through physical simulation. Only after accurately “guessing” the entire world’s physical and three-dimensional properties can we make precise predictions about it.

Therefore, we are taking a path focused on physical simulation, but we do not deny the feasibility of purely video-based models. This is our perspective during the research process.

Chen Yu: But one important point about using video models is that video data is relatively easier to obtain. Do you think the former approach is technically more difficult to implement?

Huang Xiaohuang: We also have a lot of underlying data coming from videos. But the fundamental difference lies in whether the final computation is based on video models or physical digital data—it’s not that we can’t extract information from video data. I think these two approaches share about 80% of the work, and only diverge in the final 20%.

Even within our company, there has been ongoing debate between two perspectives: since large language models can achieve remarkable results simply by scaling up—if data were infinite—why couldn’t they simulate gravity at 9.8? One side argues that gravity at 9.8 is simply a physical parameter that can be solved directly, yet you might need millions of videos to train a model to learn it.

So there are differing opinions here. Additionally, I personally have another assessment—pure video models may be a weak spot for big companies, and we’re also hesitant to pursue them.

Chen Yu: So what’s your take on this, General Manager Luo? In a sense, your company and ByteDance are also competitors.

Luo Yihang: Our approach is to first clearly define the endpoint, then work backward to determine the necessary path. First and foremost, this model must establish a closed loop from perception, understanding, and prediction to action. Second, I don’t believe a single pathway will suffice—just as in language models, beyond LLMs, and in video models, beyond DiT, there are now autoregressive approaches as well, and they are increasingly converging.

So I believe the world model will also follow an integrated path, because each single path will eventually hit a bottleneck.

Our own Motubrain follows a unified integration approach, incorporating video data, physical reinforcement data, proprioceptive data, ego data, and more.

We believe it must satisfy two conditions: first, it must be like a human—not merely taught step by step, but required to “observe,” then “think” after observing—for example, when we reach for an object, our brain already has a model of the world; it must predict, imagine, and then act. This is the essential logic we believe in; conversely, we think the architecture will continue to evolve and has not yet converged.

Chen Yu: But here’s the issue: many things are not visible. For example, when I pick up a cup with my hand, it’s very difficult to determine exactly how much force is required just by watching a video.

Luo Yihang: In pre-training, mid-training, or fine-tuning data, physical data and haptic data themselves constitute a form of data; the data involved in the subsequent RL alignment process also includes these types of data, which are inherently part of the input signals.

Chen Yu: Alright, General Manager Zhang, what do you think?

Zhang Zhizheng: World models have been the topic I’ve been asked about the most this year.

Across various industries, everyone is talking about world models, but researchers in different fields have varying definitions and understandings of them: those previously focused on content generation and video generation say that a video generation model is a world model; however, researchers in embodied AI argue that only a model capable of predicting the future and generating actions qualifies as a world model.

Actually, I believe that everyone’s current public understanding of world models is incomplete. If we return to the essence, what kind of model truly deserves to be called a world model? I summarize it into three essential capabilities, none of which can be missing.

The first capability is that it must be able to model: given the current state and a specific action you take, how does that action change the state? Whether you're predicting the future in pixel space or latent space, you must be able to model how your actions affect the environment.

The second capability is that, after modeling this change, you must be able to define a task and use the state modeling to enhance policy learning—accelerating skill acquisition and producing more reliable actions.

The third capability—one that many in the industry currently overlook—is the ability to make general evaluations. Is the change in the environment beneficial? Is the outcome of my current action what I intended? Underlying this is a general evaluation system, which is something humans inherently possess.

With these three capabilities, we can build a complete Physical AI system. All three are essential—without any one of them, skill learning cannot be sustained, and the system cannot be fully functional. Therefore, when developing embodied intelligence large models, we consider these three aspects as an integrated whole.

03 Three Capabilities of World Models

Chen Yu: Just like with large language models, the model’s inherent capabilities are certainly important, but we’ve also seen a lot of discussion lately about Agent Harnesses—systems outside the model that guide and supervise its execution and task planning. Have you observed something similar—a kind of Harness—existing outside the model in physical AI?

Zhang Zhizheng: Our team has a very clear understanding of the relationship between the Harness system and the model, because the model determines the fundamental level that your Harness system can achieve.

For example, if the Harness system is correctly感知 and well-planned but the model lags behind, it still cannot demonstrate reliable external interaction capabilities; conversely, if I only focus on the model without building a Harness system, I cannot collect more data, and the data flywheel cannot spin. Therefore, the relationship between the Harness system and the model is this: the Harness system serves as a data flywheel for the model to continuously collect data and iteratively improve its capabilities, while the model supports the Harness system, making its functions more comprehensive and reliable.

As embodied intelligence and Physical AI models, we hold them to higher standards than large language models or multimodal large models: their outputs must exhibit minimal uncertainty, closely align with the laws of the physical world, and be seamlessly integrable into a Harness system to continuously interact with the environment.

Chen Yu: So, what do everyone think—if models become increasingly powerful in the future, will Harness eventually be fully integrated into the models themselves?

Zhang Zhizheng: I believe this is an inevitable trend.

The Harness system is currently built around a model, but at its core, as a data flywheel, it uses Harness as a channel to feed back results from other models or its interactions with the environment as learning signals, thereby enhancing the capabilities of our base model.

So ultimately, my assessment of Physical AI and Digital AI is “One Model, One System”—a single model constitutes a fully functional system capable of building a complete product, or even an organization.

Chen Yu: Do you think the Digital AI and Physical AI models will eventually merge into one?

Zhang Zhizheng: I think it’s a trend. When we talk about spatial intelligence and Digital AI, these are actually subsets of the capabilities required by Physical AI—such as perception and planning, both of which are components of Physical AI. The same is true for humans: we can consume digital information on the internet while also interacting with the physical world.

So ultimately, the model will, like humans, integrate digital AI capabilities as part of physical AI into a fully functional physical AI system.

04 The Data Challenges of Physical AI and a Radical Judgment

Chen Yu: Thank you, General Manager Zhang, for your detailed response. Let’s move on to the next question related to Physical AI.

For Physical AI to further develop and be deployed, it will inevitably face many challenges, some of which are similar to those previously encountered by large language models—such as how to collect sufficient data and manage computational power between edge and cloud devices. We would like to hear from each of you: What are your views on the deployment of Physical AI, and what challenges might it face in its further development?

Huang Xiaohuang: Let me start. We don’t develop robots ourselves; instead, we collaborate closely with robot companies. Through these partnerships, we’ve realized that the future of the entire industry will be built on three pillars: data, models, and hardware.

Our efforts are primarily focused on data and models; we develop models mainly to generate data, and we also collect data from the physical world through 3D reconstruction.

Interestingly, when it comes to actual training, our computational requirements are far lower than those of large language models like DeepSeek—because we are still in the stage where data is severely insufficient.

So what we think about even more is how to collect various types of information from the physical world into the digital world.

We’ve observed that the industries where large models have been most successfully implemented are those with higher levels of digitalization, where AI delivers better results; conversely, poorer digitalization leads to weaker implementation outcomes. For example, coding has seen particularly strong results because coding is inherently digital; autonomous driving has progressed well because it automatically collects data during operation.

So we believe that to solve physical AI, the most critical step is to get as much information from the physical world into the digital world—not just visual data, but also a wide range of attributes like dimensions, mass, friction, and more. The challenge is how to collect all this information from the entire “universe” at scale, at low cost, and in bulk. Data remains the biggest bottleneck in model training, and we are currently treating it as our primary challenge. Of course, different companies may have different positioning.

Chen Yu: Manager Luo.

Luo Yihang: From the current results, the challenges are still significant. For example, it remains difficult for a long-term task involving dozens of atomic capabilities to be autonomously completed with high accuracy; if there are failure points along the way, it is also hard for the system to autonomously make dynamic decisions and still complete the task. This is from the perspective of effectiveness.

Conversely, from a systemic perspective, I see several challenges: First, the models still aren’t adequate—embodied models have not yet achieved seamless integration of perception, imagination, and decision-making, and their task completion rates and autonomous understanding remain limited; second, it’s still tied to data—issues persist with data volume, alignment between data systems, and model algorithms, commonly referred to as “non-convergence” and lack of unified approaches.

Moreover, in real-world physical systems, even more complex challenges arise: accuracy, complexity, integration with hardware, and practical implementation in commercial scenarios—all of which increase the difficulty of deploying physical AI.

Conversely, from our perspective, we also see a glimmer of potential for a solution: Motubrain’s approach to world action models is beginning to show precursors similar to the progression from GPT-2 to GPT-3 and then to GPT-3.5, in terms of data utilization efficiency and architectural differentiation—its ability to plan for long-range tasks and generalize is starting to become feasible. This is an early stage we are currently observing, but significant challenges remain.

Chen Yu: Alright, General Manager Zhang.

Zhang Zhizheng: My perspective might be slightly different—and more radical. I believe the practical application of physical AI will be broader, faster, and more far-reaching than that of digital AI.

Why say that?

This may seem counterintuitive at first, since everyone knows that building a physical AI model is more challenging. But that assumption is based on the idea that we truly aim for comprehensive generality and broad applicability—making it capable of doing and understanding everything like a human. However, the practical logic of physical AI is different; I summarize it in one sentence: strive for comprehensive generality within a limited scope.

For example, autonomous driving can be deployed as long as it achieves a certain level of generality within driving scenarios.

So today, whether we’re discussing world models, multimodal large models, embodied models, or physical AI models, we can miniaturize this “world”: for instance, confining it to the cabin where we’re having lunch today—making comprehensive generalization feel much closer. Alternatively, we can limit it to a confined space, like Galaxy General’s “capsule,” specific production lines in factories we’ve deployed, or specialized spatial skills in certain living environments—it’s finite, requiring no need to leave the room or be capable of everything.

Therefore, from this perspective, the barrier to implementing physical AI in terms of "comprehensive generality" is not as high: we can pursue its universality and generalization within a limited scope, and once achieved, we can rapidly develop a mature product along this dimension. So today, I boldly predict: the arrival of physical AI agents will significantly shorten the timeline compared to digital AI.

Chen Yu: Great, thank you to the three of you for sharing. Our boat has just docked, so we can go ashore now. Thank you all.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.