HIL-UMI: Human-in-the-Loop Post-Training for Vision-Language-Action Models

iconKuCoinFlash
Share
AI summary iconSummary
A team led by Hao Dong has developed HIL-UMI, a human-in-the-loop post-training method for vision-language-action models. The approach leverages on-chain data to identify model weaknesses during human demonstrations, without requiring robots to perform tasks. By comparing human actions with model predictions, HIL-UMI collects targeted training data to enhance performance. Tested on sorting, folding, stacking, and stamping tasks, the method improved task completion scores after three training rounds. The integration of on-chain data enables more precise model refinement.
ME AI messages: How do you know what the robot still needs to learn? A straightforward approach is to let it attempt the task first: identify where it struggles to grasp or place objects, then have a human take over to correct it. This method is highly targeted, but each batch of new experience requires the robotic arm to physically execute the actions. The team led by Dong Hao, in collaboration with Qiyuan Robotics, Peking University, and Xi’an Jiaotong University, introduced HIL-UMI, which embeds the process of identifying weaknesses directly into human demonstrations. A human operates a gripper handheld device, while the model simultaneously predicts its own actions; the system compares the differences between human and model behavior and highlights areas requiring additional training. This enables data collection and continued training based on the model’s current needs—without requiring the robot to physically perform any actions. HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface Zimu Han*1,4, Yiming Zeng*1,4, Jiyao Zhang*‡1,2,3, Zihao Zhao1, Yuanfei Wang1,2,3, Yixiang Jin5 Shiqi Li5, Shuangben Chen1, Wei Huang1, Ruodai Li5, Hui Shen5 and Hao Dong†1,2,3 Paper link: https://arxiv.org/pdf/2609.20659 Project homepage: https://hil-umi.github.io Why do errors still occur after repeated demonstrations? Suppose the goal is to put a toy on the table into a drawer. During human demonstration, the toy is typically grasped securely and placed smoothly inside. But when the robot attempts it alone, it might grasp slightly off-center, pushing the toy aside. The position and orientation change, making the next grasp fundamentally different from what was demonstrated. At this point, the robot lacks experience handling this new situation. If it continues following its original action plan, small errors can accumulate. Showing it more successful demonstrations may not necessarily address this specific gap. Currently, a class of models known as Vision-Language-Action (VLA) models can generate robotic actions based on camera input and verbal instructions. Improving existing models with new operational data is called post-training. A common method is Supervised Fine-Tuning (SFT): humans demonstrate how to complete a task, and the model learns the corresponding actions. Demonstrations help robots learn basic operations, but the conditions humans experience during successful task completion often differ from those encountered by the robot during execution. Post-training must therefore supplement data for actions the current model still struggles with. For example, if the model has learned to open a drawer but consistently places pens of different colors into incorrect slots, the next training round should focus on teaching it how to distinguish and place them correctly. If it struggles to grasp oddly shaped toys, it needs more examples of how to handle such objects. What to teach should adapt dynamically based on what the model has already learned. The Cost of Targeted Training HG-DAgger is one such targeted post-training method. It lets the robot attempt the task; when it falters, a human intervenes to demonstrate the correct action, and these correction examples are used for training. Wherever the robot performs poorly, future training focuses on filling that gap. The cost arises from this process: operators must monitor the robotic arm, wait for it to execute actions, intervene promptly, and reset the environment. To enable multiple people to collect data simultaneously, multiple robots and dedicated setups are typically required. Another tool offers a more lightweight data collection method: the Universal Manipulation Interface (UMI), a handheld gripper device that allows humans to directly manipulate objects while recording visual and motion data usable by robots. This eliminates the need for a robotic arm—but alone, it cannot tell users: Is this particular demonstration exactly what the model needs to learn right now? HIL-UMI integrates the model into this collection process. While a human demonstrates an action, the model simultaneously predicts its own action and provides feedback to guide what should be collected next. Thus, the core idea of HG-DAgger—supplementing training data based on model performance—can now be applied even when no robot execution is involved. Identifying Model Weaknesses During Human Operation During HIL-UMI data collection, a human moves objects using a handheld gripper. Cameras record visual input while the system simultaneously logs gripper position, orientation, and opening/closing state. The model runs on a computer, receiving identical visual input and task instructions to predict future actions. The human always physically moves the objects; the model’s predicted actions are never executed on a robot. The system compares human actions with model predictions to identify discrepancies. For instance, when grasping an irregularly shaped toy, a human might first adjust the gripper’s angle before finding a stable grip point. The model’s predicted approach direction or timing of closure might differ significantly from the human’s. These differences help pinpoint areas where the model still needs training. The system generates multiple action predictions; after each corresponding human action segment is completed, it comprehensively compares differences in position, gripper angle, and opening/closing state. When differences exceed a predefined threshold, the system prompts the operator to continue recording from the current state until that specific step is completed. This allows targeted collection of individual grasp or placement demonstrations—without requiring full-task recordings each time. The model begins participating in deciding what data it needs. As it updates, its discrepancies with human actions change, shifting the focus of subsequent data collection. The cycle—identifying problems, supplementing demonstrations, and retraining—can be repeated iteratively between human and model without requiring the robot to attempt each round first. Learning More Effective Actions Another challenge remains: which actions actually advance the task? Consider stacking blocks: moving a block back and forth involves motion, but aligning it precisely over another block contributes far more meaningfully to task completion. HIL-UMI uses an “advantage estimator” model to determine this. It analyzes visual input before and after an action to estimate how much progress that action contributed toward task completion. If a human action clearly advances the task but the estimator assigns it a low score, the system collects additional examples of such actions to improve its judgment. The action model then continues training by combining these advantage estimates with expert demonstrations, learning actions that better facilitate task completion. During execution, the system also provides corresponding guidance to steer the model toward generating task-advancing actions. This training approach is called Advantage-Conditioned Behavioral Cloning. Each training round uses both new demonstrations and historical data, balancing newly acquired skills with existing capabilities. The updated model then participates in the next round of data collection, continuing to provide feedback. How Much Improvement Can Three Rounds of Training Achieve with Equal Data? Can these experiences collected via handheld devices ultimately improve real robotic performance? The team tested four tasks on a Franka robotic arm using π0.5 as the base model. Desk tidying requires opening drawers, placing pens into correctly colored slots, and storing toys of varying shapes—testing sequential multi-step execution. Folding towels involves adapting to continuously changing fabric shapes. Block stacking and stamping emphasize precision: stacking requires alignment and stable placement; stamping demands precise adjustment of stamp orientation to land accurately within a designated area. Experiments were scored based on task progress—Task Progress Score (TPS), out of 100 points. For desk tidying, points were awarded for completed steps: which items were stored and which drawers were opened. Under conditions collecting an equal number of video frames, the team compared standard SFT (continuously adding generic demonstrations) with HIL-UMI (adding demonstrations based on model feedback). The chart below shows changes after three rounds of data collection and training.First, examine the purple and orange curves. Both start from the same average score of 44.3; after three rounds, SFT reaches 58.0, while HIL-UMI—with targeted data collection retained but the advantage mechanism removed—reaches 88.0. This comparison shows that whether new data aligns with the model’s weak points significantly impacts post-training effectiveness. The green curve represents the full method, which uses advantage-conditioned training from the start, increasing the average score from 62.3 to 94.8. After the third round, both stacking blocks and stamping achieve 100 points, while organizing the desk reaches 96. As new demonstrations are continuously added, performance across all four tasks improves round by round. Post-training no longer needs to revolve around the robotic arm. Eliminating repeated robotic executions and human interventions first transforms data collection speed. The team compared HIL-UMI with HG-DAgger on the desk-organizing task. HIL-UMI achieves 5.63 times the data collection efficiency of HG-DAgger, meaning it gathers more training data frames per unit time. Ultimately, HIL-UMI achieves a task progress score of 96, compared to HG-DAgger’s 91—delivering both faster collection and better task performance. The paradigm shift introduced by HIL-UMI lies in how post-training is organized: the model participates in deciding what to teach during data collection, and this feedback can be predicted without requiring actual robot execution. Collectors can supplement the model’s needed operational experience using handheld devices, reducing dependence on robotic hardware and runtime, and making post-training more accessible. The team plans next to explore distributed post-training, enabling operators at different locations to simultaneously collect demonstrations needed by the current model. Each collection point need not be equipped with a robotic arm; the gathered experience can be aggregated into a single model to support its continuous improvement. (Source: Ifnar)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.