The next frontier of AI innovation is converging on 3D generation.
From multimodal generation and world models to spatial intelligence and embodied intelligent robots...
3D is becoming the core engine of the next generation of AI.
Now, a legendary figure in the field of graphics and the hottest star company in the AI 3D space have joined forces.
Latest news:
Tong Xin has officially joined Meshy, founded by Hu Yuanming.
Hu Yuanming and the Taoist Elders Gather Together
According to the latest news, Tong Xin will serve as Chief Scientist at AI 3D company Meshy, responsible for developing Meshy’s long-term research strategy.
This may mean that two representative generations of Chinese computer graphics researchers—Hu Yuanming, founder of Meshy, and Tong Xin—will jointly advance breakthroughs in multimodal world models and real-time interactive systems.
Tong Xin previously worked at Microsoft Research Asia for 25 years, during which she led the Network Graphics group and served as a Global Research Partner.
Its research areas cover computer graphics and 3D computer vision, specifically including, but not limited to, material capture and modeling, texture synthesis, 3D geometry processing and modeling, light transport analysis and simulation, photorealistic rendering, and 3D facial animation.

△
Computer graphics, in simple terms, focuses on "how to create, manipulate, and render a three-dimensional world within a computer."
This field holds a foundational position in gaming, film and television, and industrial simulation, and is increasingly becoming the infrastructure for AI perception and expression.
After all, without understanding three dimensions, AI cannot truly comprehend or enter the physical world.
Tong Xin is one of the scholars who has been most deeply rooted and accumulated the greatest expertise in this field.
In 1999, Tong Xin joined the Microsoft China Research Institute, which had been established less than a year earlier and would later become known as the "Huangpu Military Academy of China's tech industry"—Microsoft Research Asia—right after completing his Ph.D. at Tsinghua University.
Over the past two decades, this place has produced countless pivotal figures who have shaped the landscape of computer research in China and beyond—but that’s another story.
As one of the first researchers at the Asia Research Institute, Tong Xin has spent 25 years here, progressing from researcher to Global Research Partner and Head of the Network Graphics Group.
Witnessed nearly all the important moments in China's computer graphics field.

△
During this period, Tong Xin published numerous papers in top-tier journals and conferences, with over 21,000 citations on Google Scholar.
These technical research efforts left Microsoft with many valuable assets, including the Xbox game development API, Xbox-compatible software, Windows 3D printing drivers, and the Direct3D graphics development toolkit, among others.
In addition, Tong Xin left another “legacy” with the Network Graphics Group at the Asia Research Institute: a group of professionals sufficient to form the backbone of the current Chinese-speaking computer graphics community.
Over 25 years, he collaborated with, exchanged ideas with, and mentored successive generations of outstanding researchers here.
Zhou Kun from Zhejiang University, ACM Fellow and IEEE Fellow, Director of the State Key Laboratory of CAD&CG, has collaborated for many years with Tong Xin, exploring from 3D printing and computational design to NeRF and neural rendering;
Xu Kun from Tsinghua University is a tenured associate professor in the Department of Computer Science and a National Science Fund for Distinguished Young Scholars; during his doctoral studies, he collaborated multiple times with Tong Xin on SIGGRAPH papers.
Wang Pengshuai from Peking University, Assistant Professor at the Wang Xuan Institute of Peking University, has been under Tong Xin’s mentorship since his internship, progressing to senior researcher and now stands as a leading scholar in the field of 3D geometric learning.
Those who have worked with him consistently give the same evaluation of Tong Xin:
Approachable and always on the front lines, capable of conducting rigorous discussions on highly technical details while clearly explaining complex technical issues in simple terms.
In Tong Xin's view, computer graphics is an applied computer science discipline, which is why all research problems arise from real-world needs and the emergence of new devices and data.
In addition, Tong Xin is also very humorous and witty; for example, he describes his research like this:
Besides appreciating the meals I cook, my family often jokes that I happily immerse myself in 3D teapots and rabbits on my screen, yet my work seems unable to change global affairs or improve people’s daily lives—I really don’t understand where my sense of accomplishment comes from.

△
This is the universally renowned "Tonglao" in the graphics community, possessing profound academic expertise yet radiating the warmth and curiosity of a child.
Also quite fitting.
AI for Fun
It could be said that among the graphics community, only star unicorns like Meshy can match Tong Xin’s achievements at this stage.
Meshy is the first company founded by Yuanming Hu after completing his PhD at MIT, and its mission is simple: turn text and images into 3D models.
Currently, Meshy has gathered 12 million users, serving half of the world’s top ten companies by market capitalization and valuation.
In July this year, Meshy completed a B-round financing of nearly $400 million, with a post-money valuation exceeding RMB 10 billion.
Set dual records for the largest single funding round and highest valuation in the global AI 3D sector, becoming the most highly valued company in this field.
However, whether faster growth or higher valuation, both are still too specific compared to Meshy’s vision.
Hu Yuanming, founder of Meshy.ai, 400-meter champion at Yangzhou High School sports meet, second prize winner in the Tsinghua Yao Class singing competition, Ph.D. in Computer Graphics from MIT, nominee for the SIGGRAPH Best Doctoral Dissertation, and creator of the Taichi programming language and compiler.

△
He proposed a goal that sounds a bit unusual:
AI for Fun.
His reasoning chain is as follows.
In the future, after AGI resolves a vast number of productivity issues, humanity will be left with only one core problem.
How do we create, express, and connect? Most importantly, how will we find meaning?
Reading, traveling, playing games, binge-watching short dramas... these certainly bring us joy, yet in an era where free time is so abundant, we will inevitably seek out newer, more effective ways to generate happiness and a sense of meaning.
This approach is destined to be driven by AI.
Therefore, "using AI to bring happiness and a sense of purpose to humans will be one of the most important issues over the next five years."
And Meshy aims to become the most crucial piece of this puzzle.


In Hu Yuanming’s words, what they aim to do is “redefine graphics through generative AI to find the global optimum for turning imagination into reality.”
What does this mean?
In other words, today, AI 3D must do more than simply use computer graphics to enable AI to complete modeling, texturing, and rendering more efficiently.
In the long term, the ultimate outcome of AI 3D may be the materialization of imagination.
This forward-looking perspective aligns well with Tong Xin’s philosophy.
In 2016, Tong Xin proposed addressing the "everyone and everywhere" challenge in graphical content creation, enabling anyone, anywhere, to create visual media content.
Today, ten years later, “everyone everywhere creating visual media content” has become “enabling anyone to turn an idea into an immersive, interactive world with just a single sentence.”
This is certainly a grand aspiration.
To achieve this, Meshy is currently conducting foundational technical research at three levels:
First and foremost—reinvent graphics.
To achieve this, we must move beyond the local optimum of traditional computer graphics rendering pipelines and use AI to create a new, global optimum solution that does not rely on "triangles"—the smallest unit GPUs use to render 3D objects.

△
Second, you need to set up the director system.
If the new graphics are more like a film crew, then a corresponding director system is required to achieve AI for Fun.
This director system is responsible for establishing the world's operational mechanics and 3D skeleton, enabling real-time script authoring.
Finally, infrastructure.
A low-latency, high-quality, and cost-effective 3D generation and rendering infrastructure is essential, because even a momentary lag in an AI for Fun experience will be amplified hundreds of times.
To achieve this, a new high-performance infrastructure must be redesigned, a goal that aligns closely with Hu Yuanming’s PhD thesis and the foundational work behind the Taichi programming language developed during Meshy’s early days.
Beyond these foundational research efforts, Meshy’s new model is addressing specific bottlenecks in today’s AI 3D landscape.
For example, the issue of "controllability"—whether the generated result faithfully adheres to the input image, as reflected in the accuracy of overall proportions, the reasonableness of spatial distribution, and the fidelity of surface details.
On August 10 of this year, Meshy released Meshy-7, making significant progress in geometry alignment, specifically:
In terms of organic elements, the model accurately reproduces details such as facial expressions, anatomical structure, and skin folds;
On hard surfaces and mechanical components, each part accurately lands in the designated position in the image, with adjacent parts not adhering to one another;
The engraved text and relief patterns are clear and clean, making them suitable for direct use in industrial applications such as laser engraving.
Meshy’s exploration in these areas aligns with Tong Xin’s interests, as his recent research focuses on the intersection of 3D and video generation.
As he posed in 2024, the classic question: “Is 3D merely a special case of video generation?”
That is, if a video already allows an object to be viewed from any angle, is it still necessary to explicitly construct its 3D model?
The answer is not yet clear.
What is clear right now is that Tong Xin’s cutting-edge research and engineering expertise in AI multimodal training can take Meshy’s exploration further.
Mora: Beyond the World Model
At the time this article was written, Hu Yuanming released Meshy team’s new work, Mora, on his official WeChat public account.
Mora is an acronym for "Multimodal Open-world Real-time Architecture," meaning "Multimodal Open-world Real-time Architecture."
This is quite interesting—feel free to try out the interactive game demo generated with it.
In simple terms, it consists of three pieces:
Coding agent, used to generate the skeleton and runtime code for a game world;
3D generation, which is the most crucial piece of the Meshy puzzle, can be used to enrich skeletons and generate control signals for video models;
Video model that receives 3D scene signals and outputs the final visuals and audio.
Hu Yuanming called this a technology that "goes beyond world models."
That's a bold claim—how exactly do you plan to surpass it?
Recall that in January this year, Google Genie 3 launched an interactive demo, claiming it was "a general-purpose world model that enables users to generate explorable, realistic environments using text."
As a result, Unity fell 24% in one day, and Roblox dropped 27%—the market is too excited: apparently, future game development won’t need game engines, just a video model.
However, after spending over half a year experimenting with every available world model demo on the market, Hu Yuanming outlined eight major flaws in nearly all interactive video models.
Includes weak consistency, simple interactions, limited physical computation capabilities, short experience duration, inability to support narratives and complex logic, primarily single-player gameplay, inability to leverage coding mechanisms for generation, and high latency.
In short, it's hard to play.
However, this is likely not a problem that can be solved through optimization alone.
Current video models do not possess true intelligence; they are essentially just renderers, and continuing down this path will only lead to walking simulators.
Mora’s approach is the opposite: since video models can’t solve 3D spatial consistency, let them focus on what they do best.
Have the Coding agent build the skeleton, let Meshy generate the 3D model, and finally let the video model produce the visuals—this approach ensures a stable, refined, and interactive space.

It’s clear that Boss Hu is a truly passionate veteran player.
However, Mora 1 is still in a very early stage.
It is only used to validate a framework. Then, within this framework, the target is gradually approached through scaling.
Returning to Tong Xin’s question, is 3D merely a special case of video generation?
Mora provided an initial response: at least for now, the two do not need to be substitutes; they can each fulfill their respective roles under the coordination of a Coding agent.
This answer is certainly far from definitive, but by this point, the question has already become unclear.
This feeling was not unfamiliar to Tong Xin.
Over the past two to three decades, he witnessed wave after wave of technological breakthroughs: once, creating 3D graphics relied entirely on handwritten rules; then neural rendering emerged, enabling AI to learn lighting and shadows from photographs; with generative AI, photorealistic imagery no longer depended on traditional 3D pipelines; and now, video models have arrived, capable of generating dynamic worlds from nothing.
New technologies continually challenge the boundaries of traditional graphics: how should graphics continue to exist?
Faced with this increasingly urgent issue, Tong Xin set out once again.
He wants to create a new graphics with the most imaginative generation of young people.
This article is from the WeChat official account "Quantum Bit," authored by Cheng Qian.
