Global First: Guanglun Intelligence Opens Source 100,000 Hours of Human Data

iconMetaEra
Share
AI summary iconSummary
Guanglun Intelligence unveiled EgoSuite-Open100K at the 2026 World Robotics Conference—a 100,000-hour first-person human behavior dataset open-sourced on Hugging Face. The dataset encompasses 15,000+ real-world scenes across seven environments and 128 categories, featuring detailed annotations for hand and body poses and semantic events. It enables robot training at scales of 10,000 to 100,000 hours. Hugging Face’s LeRobot project commended the release, labeling it a milestone for embodied AI. This launch occurs amid evolving global crypto policies and growing demand for high-quality training data, particularly as inflation data influences investment in AI and blockchain infrastructure.
Lightwheel Intelligence announced at the 2026 World Robot Conference the open-sourcing of EgoSuite-Open100K, a 100,000-hour fully annotated first-person human behavior dataset—the largest of its kind globally.

Article author and source: AI New Era

Last week, a company called Dyna Robotics did something that kept the entire industry awake.

They trained a model called DYNA-2. The unique aspect is that no robot action data was fed during the pretraining phase.

Instead, they fed the model over one million hours of human first-person video—equivalent to a person learning continuously for 170 years.

As a result, from 1,000 hours all the way to 1 million hours, the robot becomes smarter with every tenfold increase. More precisely, it follows a power-law trend that continuously rises without any noticeable plateau.

It can be said that physical AI finally has its own scaling law!

As soon as the news broke, the industry was shaken—because it essentially declared that the most scarce resource for training bots isn't GPUs or algorithms, but high-quality video data of human behavior.

Whoever holds this data controls the oil of the physical AI era.

But the problem arises: while this curve is visible, the data needed to reach the critical turning point is locked away on each company’s own hard drives.

Until the company broke the lock.

On the morning of August 20, at the 2026 World Robot Conference, LightWheel Intelligence announced: 100,000 hours of multimodal human behavior data is now fully open-sourced!

The dataset is called EgoSuite-Open100K, jointly released with Hugging Face, and is now available on both the Hugging Face Hub and AtomGit. It can be used for academic research and commercial training.

The world's largest fully annotated open-source first-person human dataset—the most valuable asset for embodied AI—is now available to everyone for the first time.

After the message was posted, it quickly blew up on X.

Hugging Face itself was the first to explode. The official HF account directly reposted the announcement from LightWheel AI.

Anyone familiar with the world’s largest open-source AI community knows that the Hugging Face official account rarely publicly endorses any third-party project on its own.

A search through HF’s official account history reveals only a handful of similar actions: the Surya solar physics model released jointly by IBM and NASA; Databricks’ CommonCanvas text-and-image dataset, which the official account called “the largest dataset ever hosted on Hub”; and the image generation model Flux.2-dev. Each represents a landmark open-source event in its respective field.

EgoSuite-Open100K received the same treatment as others at its level.

Not just the official account. HF’s robotic project, LeRobot, also posted a lengthy message in strong support, explicitly stating: “Always great to see a dataset at this scale released openly on the Hub.”

Following up, prominent robotics influencer Lukas Ziegler provided a detailed analysis of the dataset's multi-view variants, calling it "an invaluable resource for training dexterous manipulation."

Assistant Professor Jiafei Duan of the National University of Singapore said: Open source is the new moat.

Why specifically 100,000 hours?

Behind this number are a series of rigorous research findings in the field of physical AI over the past year.

NVIDIA's EgoScale framework, released in 2026, trained a VLA model on 20,854 hours of first-person human video, for the first time establishing a clear log-linear scaling curve in embodied intelligence—where the number of human data hours correlates log-linearly with validation loss (R² = 0.998), achieving a 54% average success rate improvement on a 22-degree-of-freedom dexterous hand.

The scaling laws for physical AI are becoming increasingly clear: the more human demonstrations and the broader the range of real-world scenarios covered, the stronger the robot's action model becomes; more importantly, this improvement is now beginning to be measurable and predictable.

The generalist AI trained GEN-0 using 270,000 hours of physical interaction data. Sunday Robotics (the team behind ALOHA) demonstrated that human demonstration data collection can be scaled beyond the lab and operated effectively in real-world home environments.

Then there's Dyna Robotics, mentioned earlier. The most important insight from DYNA-2 isn't just that "human videos can train robots."

More importantly, they pinpointed the tipping point for cross-embodiment transfer to be between 10,000 and 100,000 hours. In simple terms, if you show a six-axis robotic arm videos of humans chopping vegetables and train it, after training, it can chop vegetables just as well on a humanoid robot.

This essentially never happens in under 10,000 hours. If you teach it to chop carrots with knife B in kitchen A, it will only do so in kitchen A with knife B—even changing the table won’t work.

But after 10,000 hours, the curve begins to bend. At the level of 100,000 hours, the signals for embodied transfer become clear and stable. In other words, the range from 10,000 to 100,000 hours is the threshold where robots "wake up."

Ironically, most teams still have far too little data. This has created a situation where everyone knows about the Scaling Law, but only a handful of players are qualified to test it.

Guanglun Intelligence's release of this 100,000-hour dataset has arrived precisely at this turning point.

How to create a high-quality "encyclopedia of instruction and example" for a bot?

100,000 hours of data isn’t something you can accumulate just by leaving the camera on.

In data collection, scale, diversity, and annotation quality form an impossible triangle.

When you scale up, your reach narrows. If you hire ten thousand people but they’re all chopping vegetables in the same kitchen, 100,000 hours makes no difference from 10,000 hours.

The coverage has expanded, but the labeling depth can't keep up. The more complex the scenario, the more elements need to be labeled per frame, causing labeling costs to rise exponentially. If you label in sufficient detail, scalability suffers because you can only label a few samples per day.

Guanglun Intelligence designs diversity into the data collection process from the start. It’s not feasible to shoot a large volume first and then filter later—there’s no time, and the cost of waste would be unsustainable.

They operate according to the same set of standards. Regardless of the scenario or task, every step is strictly controlled to meet the target coverage.

The final collected data includes: over 15,000 unique collection scenarios (each representing a real physical space with distinct layouts, objects, and lighting), over 15,000 specific tasks, seven major environmental categories—residential, hospitality, retail, sports, logistics, office, and industrial—128 scene types, and 18 task categories.

Covers real-world tasks ranging from kitchen food preparation to industrial assembly, from inventory counting to tool maintenance. Currently available public first-person data is overwhelmingly concentrated in home kitchens and daily activities, with only a handful of independent scenarios. EgoSuite-Open100K directly expands the scope of scenarios by an order of magnitude.

Most importantly, these are precisely the areas where the robots will actually be working, and they represent the largest gaps in existing public datasets.

One hundred thousand hours represents quantity, but what truly determines the value of this dataset is quality. Currently, only this platform offers the combination of the world’s largest open-source first-person dataset, the most comprehensive three-layer annotation, and the broadest scene coverage.

Each video in EgoSuite-Open100K comes with three layers of annotations:

  • Human hand poses teach the robot how to bend and pinch its finger joints—21 keypoints, frame-by-frame annotated;
  • Teach it body posture: how to reach out, how to bend over, and how to coordinate the entire body;
  • Event-level semantic annotation tells it what this step is doing and which object is being manipulated.

Three layers stacked together enable the model to understand both "how the hand moves" and "why it moves that way."

The entire dataset is divided into two streams. EgoStandard is the primary stream, capturing first-person perspective footage aligned with the user's eye movements.

EgoPro additionally features a wrist-mounted camera that captures the precise moments when fingers make contact with objects—the angle at which fingers touch screws, the force applied to grip bag openings, and the timing of inserting cards into slots—all of these fine details are obscured by the arm in the head-mounted perspective.

The wrist-angle view fills this gap—it's expensive to collect and almost never found in public datasets.

The most challenging part is the annotation itself—and this is precisely what makes Lightwheel AI’s open-sourced dataset the most “genuine.” At this scale, similar datasets typically release only raw video footage, as annotation costs are prohibitively high.

The head-mounted video keeps shaking—晃动 during walking, turning the head, and squatting or standing up. The hand joints have high degrees of freedom and are frequently occluded by the body, objects, or each other, often precisely during the most critical frames for operation.

Occlusion combined with jitter is a nightmare for pose estimation and represents the area where Lightwheel Intelligence has invested the most engineering effort in its annotation pipeline.

Therefore, this dataset, which includes multimodal information, is of significantly higher quality than datasets containing only first-person data.

The biggest issue in the embodied AI industry today isn't just insufficient data—it's that the data collected by different companies simply can't be combined or used together.

The data collection methods differ, the labeling standards differ, the data formats differ, and the temporal organization differs. If you collected 10,000 hours of data using Device A and I collected 20,000 hours using Device B, combining these two datasets would leave the model unable to process them.

The opening of EgoSuite-Open100K first addresses this issue: enabling everyone to speak from the same data foundation. It is only natural that Guanglun Intelligence initiated this effort, as it is a member of the EgoVerse Global Human Data Council and, alongside leading institutions such as Stanford and ScaleAI, is collaboratively building open standards for human data to transform it into an “encyclopedia of instruction and example” for continuous robot learning.

For researchers, this is the first time there has been a publicly available, fully annotated, and reproducible human data benchmark. In the past, conclusions in papers often “talked past each other” due to small datasets and missing annotations; now, there is finally a global coordinate system everyone can align with.

From Demonstration to Application: How to Validate Data Value

After reading the encyclopedia, has the robot truly learned?

Human videos provide behavioral priors, while simulations enable scalable trial-and-error and evaluation.

Lightwheel Intelligence's fully self-developed SimFoundry simulation platform is the technological foundation that enables robots to repeatedly train, experiment, and receive corrections after reading an encyclopedia.

First, extract tasks, action sequences, hand poses, and event boundaries from human demonstrations to form reproducible task definitions; then map real objects, materials, contact relationships, and task logic into a simulation environment, repeatedly training under varying scenarios, object states, and physical perturbations; the system records the model’s task success rate, trajectory deviation, completion time, collisions, and failure types; finally, use real-world playback to calibrate simulation results and verify whether improvements observed in simulation translate to real-world tasks.

Thus, human data addresses the question of “what robots should learn from humans,” simulation addresses whether such learning can be scaled and reinforced, and large-scale evaluation further answers “how much the model has actually improved.”

The scale evaluation is RoboFinals. The logic is simple: after robots finish learning, they must take an exam—you won’t know if they’ve truly learned without one.

The industrialization of embodied intelligence must answer three questions simultaneously: Can it be mass-produced? Can it perform tasks? Can it operate stably?

Embodied intelligence is a complex system. The hardware, models, environmental conditions, and task workflows are all intertwined—changing any single variable can lead to entirely different outcomes.

So the evaluation criteria have also changed.

Don’t just look at whether tasks are completed—also understand: At what success rate can tasks be consistently completed? What conditions cause performance degradation? Where is the degradation boundary? What are the causes of failure? Has performance regressed after version updates? How large is the discrepancy between simulation and real-world conditions?

RoboFinals is here to answer these questions.

It is powered by Lightwheel Intelligence's proprietary SimFoundry physics simulation platform, supporting thousands of instances running simultaneously with time unaffected by task volume.

By standardizing task definitions and fixing random seeds, results produced by different teams at different times can be directly compared. This is crucial, as most evaluation results in the industry today are inconsistent and cannot be meaningfully compared.

Real-device testing takes 2–3 weeks per round and costs over 500,000, with additional security risks—deploying untrained bots on real devices may damage equipment, but worse, could cause injury.

RoboFinals moves this process into simulation, compressing each test cycle to the hour level and significantly reducing costs—no matter how many times you test, it’s all virtual, with no risk to real machines or people, enabling scalable evaluation across multiple scenarios, agents, and models.

Conclusion

Behind the world's first open-source dataset of human behavior with over 100,000 hours of multimodal data, Guanglun Intelligence is telling a much more complete story:

Human data is the textbook, simulation platforms are the practice fields, evaluation systems are the exams, and real-world deployment is the job—when you fail on the job, go back and update the textbook. This is a continuous learning loop of Real-to-Sim-to-Real.

Another notable point is Guanglun Intelligence's positioning according to international standards.

It is not only a member of the EgoVerse Human Data Committee, collaborating with Stanford, Scale, and others to establish human data standards; it is also a member of the Newton Simulation Technology Committee, working with NVIDIA, Google DeepMind, Disney Research, and Toyota Research Institute to set simulation standards. Additionally, note that Guanglun Intelligence is the only Chinese company among these two international embodied standards committees.

EgoSuite was built following these guidelines, so its data can integrate with the work of other members rather than becoming another isolated data silo.

Simulation standards determine "how to learn," while data standards determine "who to learn from." Among Chinese companies seated at both of these tables, only this one exists.

Over the past few years, the most valuable assets in embodied intelligence have consistently been these human behavior datasets—data that no one is willing to make public. They determine what robots can learn, how deeply they can learn it, and who in this industry can truly get a seat at the table.

Today, 100,000 hours of open educational materials have been made available.

But 100,000 hours is just the beginning. Lightwheel has also announced a larger initiative: 10 billion hours of embodied intelligence data collaboration over five years. To achieve this, they’ve outlined three key actions—unifying data standards across different devices to enable cross-device and cross-task data reuse; identifying data gaps from model failures and conducting targeted data collection and generation, driven by evaluation to create a closed data loop; and co-designing data and models to optimize the combination of simulation data, human data, teleoperation data, and reinforcement learning data.

In addition to establishing a closed loop for continuous learning, we must also recruit industry partners to co-build an open ecosystem for ongoing learning.

As Dr. Xie Chen, Founder and CEO of Guanglun Intelligence, stated at the WRC main forum:

Human civilization stems from collective intelligence; robot civilization should not be created by a single company either.

The leap from 100,000 to 10 billion—five orders of magnitude—will become a new civilization co-created by humans and robots.

When ASI steps from the screen into the physical world, it must first understand how humans pick up a cup, how they feel for a screw despite obstructions, and how they correct their own mistakes.

Humans have practiced these movements for millions of years, embedding them into muscle and instinct. For machines to learn the same things, they must watch hour by hour and frame by frame.

The so-called 10 billion hours essentially means translating hundreds of thousands of years of human physical experience into a language machines can understand.

There are no shortcuts to this work.

One hundred thousand hours is the first page of this textbook.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.