[Axis’s Public Release of 100 Hours of Egocentric Data: How to Convert Human Actions into Robot Learning Data] @KaitoAI X @axisrobotics Axis Participation Link (Referral): https://t.co/hUbWSVsBVR When humans pick up a cup, they don’t calculate finger angles or wrist trajectories—they simply reach out and naturally grasp the cup in front of them, then place it elsewhere. Teaching a robot to do the same requires a different approach. You must break down the action into distinct phases: when the hand approaches the object, the moment of contact, the path taken while lifting and moving it, and the point of placement. What humans perceive as a single fluid motion must be decomposed into data points for a robot to learn. Axis’s updated axis-ego-samples, released in early September, illustrates this transformation of human behavior into robot learning data through actual files. The publicly released egocentric video now totals 100 hours. These first-person videos, captured from the worker’s viewpoint, are categorized by real-world work environments such as education, hospitality, retail, warehousing, craftsmanship, catering, office work, manufacturing, and logistics. The data is organized in LeRobot format and includes primary dense action annotations within the same dataset. The actions shown are unremarkable: retrieving objects and placing them elsewhere, using tools, or tidying workspaces—everyday scenes commonly seen in homes or workplaces. Yet for these ordinary actions to become useful robot learning data, they must be annotated with far greater precision. The annotation review contains 101 clips, segmented into 80,090 frame-by-frame language annotations. Rather than assigning a single task label to an entire long video, each segment is labeled precisely at the moment an action changes—documenting exactly what motion occurred at each instant. Even the simple act of moving a cup is broken down this way: the phase when the hand approaches the cup, the exact moment of grasping, the state during movement, and the point of placement are all separately identified. What appears to the human eye as one natural motion becomes multiple distinct steps in the data. This level of detail is necessary because human and robot bodies are fundamentally different. Human hands freely coordinate multiple joints and fingers, while robots vary widely in structure and range of motion. Some use simple grippers; others have cameras mounted in different positions. You cannot directly copy human hand motions into robot actions. Instead, you must first identify action segments and object states in the video, then map them to action representations that the specific robot can physically execute. Multiple types of capture devices are used. The axis-ego-samples release includes samples collected from diverse hardware: dual fisheye cameras, six-camera RGB+SLAM systems, MEgo, Gen DAS Ego, FEAGINE EGO, GI Labs, and others. Even when recording the same task, different camera setups yield different visual inputs: varying fields of view, differing hand positions within the frame, and distinct methods for capturing spatial information. When collecting large-scale real-world data, these variations are unavoidable. Rather than assuming uniform equipment or conditions, you must also prepare the data to accommodate diverse input sources for robust learning. A separate folder, sample-150h-10cat, contains 10 episodes selected from a 150-hour delivery corpus, categorized by task type. However, only these 10 samples are publicly available—the full 150-hour corpus has not been released. The total file size of the entire axis-ego-samples repository is approximately 770 GB. You can also see how this data connects to other Axis products. The Mobile Egocentric App is designed to track hand poses in real time as a person moves and translate those motions into robotic actions. In browser-based simulations, users directly control virtual robots to generate state-action pairs. In egocentric capture, real-world human activities are first recorded, then analyzed to extract robot-relevant actions. One approach involves generating data by directly controlling robots; the other transforms existing human behavior into training data. Here lies an interesting insight: To give robots experience, you don’t necessarily need to keep moving robots continuously. Humans already handle countless objects every day—picking up cups, opening doors, organizing boxes, using tools. These actions occur constantly in everyday life. If these behaviors can be accurately recorded and mapped to robotic motion representations, then human daily activity itself becomes a rich source of robot learning data. How this concept is applied in actual robotics research can also be observed in EgoScale—a study independent of Axis.EgoScale, unveiled in February 2026, first pre-trained a VLA on over 20,854 hours of action-labeled egocentric human video, then further fine-tuned it on human-robot aligned data. On the 22-DoF dexterous robotic hand evaluation, the final policy achieved an average success rate 54% higher than the no-pretraining baseline. EgoScale is not an Axis-related project, nor does it imply that simply feeding large amounts of human video enables robots to move well immediately. It demonstrates that performance improves when robots first learn broad experience from human behavior, then receive additional data tailored to their own embodiment and actions. Axis is actively connecting this data to real-world collaboration. In a collaboration announced on August 28 with Dexmal, Axis revealed it is responsible for producing large-scale egocentric, simulation, and real-world data to be used in Dexmal’s VLA and world model. Some details remain undisclosed: the exact volume of data delivered to Dexmal and the specific impact of this data on model performance have not been revealed. Even with only the publicly available information, it is clear what kind of work Axis is undertaking. 100 hours of LeRobot data have been publicly released, with frame-by-frame annotations segmented into 80,090 distinct clips. Samples captured using multiple devices are also included. The process involves filming short, natural human actions performed in real life, segmenting the boundaries of those actions, annotating object and hand movements, and converting them into a format suitable for robot learning. What humans casually experience in a few seconds becomes a single learning experience for robots. What Axis is now making public is the tangible result of this process, captured in real-world data. #PhysicalAI #RobotLearning
Grid (❖,❖) 🟩 🐬TermMax 🚢Share

Source:Show original
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information.
Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.