Being-M0.7: The World’s First Implicit World-Action Model for Full-Body Mobility in Humanoid Robots

icon MarsBit
Share
AI summary iconSummary
Being-M0.7, the world’s first implicit world-action model for full-body mobility in humanoid robots, is now live. Developed by BeingBeyond, the model leverages over 10,000 hours of human-centered multimodal data and handles complex tasks such as fish retrieval and obstacle avoidance. The Vision-Motion MoT architecture unifies visual and motion data for efficient training. As news around real-world assets (RWA) continues to grow, such innovations align with global cryptocurrency policy shifts.

Over the past two years, the competitive focus in the humanoid robotics sector has expanded further from whole-machine hardware to model capabilities.

Major manufacturers are rapidly launching new products, with videos of robots doing backflips, dancing, and running marathons dominating social feeds. Yet behind the hype, the industry is increasingly converging on a consensus: the upper limit of humanoid robot capabilities is no longer determined solely by joints and motors. The ability to understand the environment, anticipate changes, and coordinate the entire body to accomplish tasks is becoming key to their path toward general-purpose functionality.

World models, VLA, and humanoid robot foundation models have thus become among the most important technical directions in this field over the past two years.

Beneath the excitement, three major challenges continue to face the entire industry.

First, the cost of collecting real-world demonstration data for humanoid robots is high; during data collection, it is necessary to simultaneously record first-person video, proprioceptive data, and executable full-body commands. Due to limitations in teleoperation difficulty, safety risks, hardware availability, and environmental diversity, it is challenging to accumulate large-scale, high-quality data in a short time.

Second, many existing world action models rely on pixel-level video prediction, which incurs high computational costs and wastes significant capacity on visual details weakly related to control. The rapid self-motion and camera jitter of humanoid robots further amplify visual prediction noise.

Third, many existing approaches model upper-body actions and locomotion control separately, resulting in insufficient coordination between the upper and lower body and making it difficult to achieve natural, fluid full-body control.

Under this context, embodied AI company Zhi Zai Wu Jie has launched Being-M0.7—the world’s first Latent World-Action Model (Latent WAM) designed for full-body mobility and manipulation in humanoid robots, and the first in the industry to extend implicit world model capabilities from dexterous tabletop manipulation to full-body mobile manipulation.

Wisdom Without Bounds

  • Paper link: https://research.beingbeyond.com/being-m07/being-m07.pdf
  • Project homepage: https://research.beingbeyond.com/being-m07

It was pre-trained on over 10,000 hours of human-centered, multimodal data, fine-tuned with a small amount of real-world demonstration data for embodiment adaptation, and successfully performed multiple challenging whole-body mobility tasks on actual humanoid robots.

From Being-H to Being-M: a consistent roadmap for delivery

Behind Being-M0.7 is a technological path that Zhi Zai Wu Jie has consistently pursued for many years.

This company is among the earliest embodied AI enterprises to bet on large-scale human video training, while also developing two core model lines: general dexterous manipulation and general mobile dexterous manipulation. It is also the first team in China to launch a native embodied implicit world action model.

The core insight of this approach is that real-world robotic demonstration data is expensive and scarce, making it difficult to scale continuously like internet text and video. In contrast, humans interact daily with the physical world from a first-person perspective, generating vast amounts of behavioral data rich with prior knowledge about scene evolution, object dynamics, and bodily coordination. Rather than waiting for robotic data to accumulate slowly, it is more effective to first train models to learn how the world works from human experience, then transfer this knowledge to specific robotic platforms.

Being-H0.7, released in April this year, validated the feasibility of this approach on the dexterous manipulation side. The model scaled its training data to 200,000 hours of human video, achieving the top overall ranking globally across six international benchmarks, with first-place finishes in four of them. It became the first general embodied world model to cover seven key dimensions: cross-ontology, cross-scenario, continuous dynamics, fluids, deformable objects, physical laws, and contextual reasoning.

Wisdom Without Bounds

Being-M0.7 is the latest achievement in this implicit world action model pathway.

While the Being-H series addresses how the hands interact with the world, the Being-M0.7 addresses how the entire body coordinates movement and action within the world. Humanoid robot locomotion and manipulation (loco-manipulation) require the model to simultaneously decide where to go, how to orient the body, how to coordinate the limbs, and how to maintain stable posture—a problem that is highly coupled across both temporal and bodily dimensions, and an essential capability that no general-purpose humanoid robot can bypass.

Wisdom Without Bounds

Being-M0.7 is an implicit world action model pre-trained on first-person video and human motion data using a Mixture of Transformers (MoT) architecture, followed by fine-tuning with action experts on robotic trajectory data from diverse whole-body manipulation tasks to achieve control deployment.

Unlike many world models that rely on pixel-level video generation, Being-M0.7 predicts future environmental states in latent space, coupling them with compact whole-body motion representations. Pixel-level prediction is computationally expensive and wastes significant resources on appearance details irrelevant to control. Additionally, intense self-motion and camera jitter in first-person views introduce substantial noise into predictions. Latent space prediction focuses modeling capacity on semantically meaningful states, object layouts, and scene evolution—the very structures directly relevant to control—preserving the core ability of world models to anticipate the future while significantly reducing computational overhead.

Physical understanding—how does it translate into full-body action?

Whether the model truly possesses full-body movement capabilities must ultimately be tested in real-world scenarios.

Zhi Zai Wu Jie has unveiled four real-device demos centered around Being-M0.7, covering four highly challenging scenarios: fluid interaction, mirror reasoning, long-range tasks, and obstacle avoidance.

These tasks collectively test whether a robot can continuously generate full-body movements that match the current scenario, based on predictions of the environment and future changes.

Netting fish from an aquarium

The robot moves to the tank and uses a handheld net to scoop up toy fish in the water. The liquid has no fixed shape, flows freely, exerts buoyancy and drag on submerged objects, and causes visual displacement of underwater targets due to refraction at the surface. The robot must understand the interactions between the water, the net, and the fish, and coordinate its arm to perform a dynamic, tool-assisted capture despite visual distortions caused by the water. This task tests the robot’s ability to predict future states, use tools, and coordinate actions under uncertain object dynamics.

For the fishing task, Being-M0.7 succeeded 3 times out of 5 tests. In comparison,

Wisdom Without Bounds

For 2/5, GR00T-N1.6 is 1/5.

Wisdom Without Bounds

Mirror item retrieval

In front of the robot is a box that opens only at the back and sides; the contents inside are completely invisible from the robot’s own perspective. The only clue comes from the reflection in a mirror ahead. The robot must infer the three-dimensional position of the hidden object based on the mirror’s reflection, then move toward the box and reach out to grasp it. This requires the model to understand the spatial relationships among the mirror, the box, and the object, as well as the principles of mirror reflection, and to transform indirect visual evidence into actionable movements under partial observability.

In conjunction with

Wisdom Without Bounds

In real-world comparative tests of GR00T-N1.6, under distance settings of 0.5 meters and 1 meter, Being-M0.7 succeeded 3 times and 1 time out of five tests each, for a total of 4 out of 10.

Wisdom Without Bounds

The overall scores for GR00T are also 1/10.

Wisdom Without Bounds

This result indicates that Being-M0.7 demonstrates greater adaptability in partially observable tasks requiring indirect visual reasoning, full-body approach, and fine grasping.

Move and retrieve items

The robot walks to the table, transfers a baguette from one basket to another, then picks up a bouquet of flowers and turns to leave. The task consists of multiple chained subtasks, requiring the robot to continuously switch between behaviors such as walking, grasping, transferring, and turning, while maintaining ongoing understanding of the scene. It tests not only the success rate of individual grasps, but also state retention over long-range tasks, object-level spatial reasoning, and full-body coordination between locomotion and dexterous manipulation.

Box avoidance

The robot moves forward while carrying a box; when it encounters an obstacle, instead of stopping completely to replan its path, it adjusts its body orientation and sidesteps through the narrow gap between obstacles. The carried object partially obstructs the first-person view and alters the robot’s load distribution and center of gravity. The model must integrate existing environmental information with real-time feedback to identify passable areas, adjust its direction and full-body posture, and maintain both its own balance and the stability of the carried object. Multi-directional movement, obstacle avoidance, and load sensing are unified into a single closed-loop behavior.

These demonstrations show that the robot does not execute open-loop movements along a fixed trajectory, but rather continuously generates and adjusts full-body motions based on current observations, real-time feedback, and predictions of the future.

MoT architecture and unified motion representation to solve the challenge of scarce embodied data

The capabilities described above are supported by a set of key design choices in Being-M0.7 at the data and architecture levels.

Humanoid robots require first-person visual information to achieve spatial awareness and must generate future movement and control commands. High-quality human motion data typically requires motion capture equipment, and paired first-person data with precise alignment between vision and motion is even more scarce.

If a model can only use data that simultaneously includes visual and motion information, the scale of trainable data will be severely limited. Being-M0.7 addresses the challenge of enabling paired data, pure video data, and pure motion data to jointly participate in training.

Zhi Zai Wu Jie chose the Vision-Motion MoT (Mixture-of-Transformers, a multimodal Transformer architecture). Vision-Motion MoT preserves modality-specific projections and processing modules for vision and motion, while enabling cross-modal interaction through shared multimodal attention. Visual state changes and continuous motion have different data distributions and do not need to be forced into an identical parameter framework; when both modalities are present, they can exchange information within a shared context.

This enables the model to simultaneously support three types of data.

For video-motion paired data, the model jointly learns future environmental states and motion trajectories; for pure video data, only the visual branch's training objective is computed; for pure motion data, only the motion branch is trained. Data from different sources jointly constrain the model through a single training objective, eliminating the need to train multiple isolated unimodal systems separately.

From a probabilistic modeling perspective, paired data captures the joint relationship between vision and motion, while unimodal data provides marginal constraints on this joint distribution. Even when data modalities are incomplete, they can still be incorporated into the same training framework.

Wisdom Without Bounds

Overview of the Being-M0.7 Training Framework. Top left: Pre-training data consists of video-motion paired data, pure video data, and pure motion data. Bottom left: The research team developed a unified motion representation shared by humans and humanoid robots, providing richer supervision signals and feedback for training and inference. Right: Overall model architecture of Being-M0.7.

Based on this architecture, the team constructed over 10,000 hours of multimodal pretraining data, including first-person human videos, paired first-person video and motion data, and pure human motion sequences.

Wisdom Without Bounds

The data recipe for Being-M0.7. The pretraining corpus is sourced from multiple external public datasets, including Ego4D, Xperience, Nymeria, Bones-SEED, SnapMoGen, HumanML3D, and Lafan1, as well as internal datasets.

Another key design is a unified motion representation shared between humans and humanoid robots.

Being-M0.7 proposes a unified action representation that converts human motion data from diverse sources into a standardized representation rooted at the head, naturally aligning with first-person vision. Through standardization steps such as unifying the coordinate system and eliminating initial orientation, it reduces distribution differences between datasets and enhances consistency across data sources.

Furthermore, Being-M0.7 employs a compact motion representation that retains only the head, hands, and feet, effectively bridging the morphological gap between humans and robots while preserving critical interaction and contact information. This representation not only provides richer supervisory signals for robot post-training than action labels alone, but also enables motion-level feedback during inference to support whole-body coordinated control.

During the pre-training phase, the model maps images to a latent space using a visual encoder and employs a compact, unified motion representation. The model is trained with a flow matching objective to jointly predict future state changes and motion trajectories based on a short history of visual-motion data and task instructions.

During the real-world data collection phase, the team built a full-body teleoperation system based on PICO VR. Operators wear a PICO headset, two ankle trackers, and two handheld controllers; the VR system estimates human posture in real time and converts it into 29-degree-of-freedom full-body control commands executable by Unitree G1. While the robot performs the teleoperated actions, it records first-person-view images from its onboard RGB camera, proprioceptive sensing data, and motion control commands as fine-tuning data for Being-M0.7 on specific tasks.

Wisdom Without Bounds

Being-M0.7 Real-World Data Collection System. Operators provide full-body motion commands via VR devices; the system converts human posture into robot control commands while simultaneously capturing first-person imagery, proprioceptive data, and motion trajectories.

Since the model has already established visual-motor priors during pre-training, real-world data no longer needs to teach all motion dynamics from scratch; instead, it primarily performs two tasks: first, adapting the pre-trained priors to the specific control space of humanoid robots; second, learning the low-level control commands and feedback mechanisms required by the actual robot. This process is carried out by a lightweight Action Expert. The Action Expert reads the intermediate hidden states of Latent WAM as high-level planning context, then combines them with current visual observations, proprioceptive information, and execution progress to generate action blocks that the robot can directly execute.

During inference, the model generates future video-motion plans at a lower frequency and converts intermediate hidden states into a reusable policy cache (KV Cache). The unified motion representation integrates both visual and proprioceptive feedback, while also correcting predicted full-body motion plans using the robot’s latest motion state, enabling the policy to promptly respond to deviations in body and end-effector movements. The action expert reuses the current KV Cache to continuously generate actions at a higher frequency, seamlessly incorporating the latest robot feedback upon cache refresh. This design decouples low-frequency world planning from high-frequency action control, ensuring real-time performance while keeping the robot guided by both long-term planning and real-time feedback.

Scalable fusion paradigm toward more general embodied intelligence

The significance of the Vision-Motion MoT architecture lies not only in solving the training challenges of the Being-M0.7 model, but in establishing a scalable, multimodal fusion paradigm.

The most direct change from this paradigm occurs at the data level.

Over 10,000 hours of multimodal data have expanded the sources of supervised signals for training humanoid robot models from expensive and scarce real-robot demonstrations to vast amounts of human behavioral data. The relaxation of this data bottleneck is a prerequisite for any scaling law to hold.

Meanwhile, Being-M0.7 has also adjusted the order of model learning.

Before being adapted into robot-executable instructions, the model first learns visual context, future dynamics, and human kinematic structures from large-scale human-centered data. Subsequently, action experts transform these predictions into specific robotic control commands by integrating them with motion priors. In other words, the model first develops the ability to predict future states and body movements, and then learns how to act on a specific embodiment. This constitutes a key distinction from traditional imitation learning approaches that directly learn action mappings from robot demonstrations—these typically begin with “see something, output an action,” whereas Being-M0.7 introduces an additional layer of joint modeling of future states and motion prior to action generation.

More importantly, this architecture does not require all new data to have complete visual-motor pairings. After cleaning and processing, individual human videos and motion sequences can both be incorporated into the same model. As the data scale continues to grow, this integrated paradigm is also expected to continually expand its capabilities.

Within the broader industry context, the release of Being-M0.7 may signify a shift in the competitive dynamics of humanoid robotics.

Over the past few years, industry attention has focused more on whose ontology is more flexible and whose motion demonstrations are more impressive. As hardware performance continues to improve, the ability of models to understand scenes, predict changes, and generate coordinated full-body movements—and whether there exists a scalable data infrastructure behind them—will increasingly become the key differentiators.

The development of large language models has demonstrated that scalable data and training feedback loops often determine how far a technological path can go. Embodied intelligence is now at a similar tipping point: unlike internet text corpora, real-world robotic data cannot grow rapidly—where else can robots obtain the experiences needed for continuous evolution?

From Being-H to Being-M, the judgment of Being-Unbounded is for the robot to first learn about the world from human behavior and then translate that knowledge into actions within the physical world.

When understanding becomes a prerequisite for action, general-purpose humanoid robots truly step out of laboratory narratives and begin to enter the physical world of countless industries.

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Yang Wen.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.