spot_img
HomeResearch & DevelopmentHuman-Assisted Online Learning for Robust Robotic Manipulation

Human-Assisted Online Learning for Robust Robotic Manipulation

TLDR: Hi-ORS (Human-in-the-loop Online Rejection Sampling) is a novel post-training method for Vision-Language-Action (VLA) models in robotics. It addresses the instability of Reinforcement Learning and limitations of Imitation Learning by using outcome-based rejection sampling to filter out unsuccessful robot attempts and providing dense, reward-weighted supervision. Crucially, it integrates flexible human interventions to guide error recovery behaviors. This approach enables robots to learn complex, contact-rich manipulation tasks more stably and efficiently in real-world settings, demonstrating superior performance and test-time scalability compared to existing baselines.

In the rapidly evolving world of robotics, Vision-Language-Action (VLA) models are becoming increasingly crucial for enabling robots to perform complex manipulation tasks. These models, often pre-trained on vast datasets, still require significant fine-tuning for real-world deployment. However, traditional methods like Reinforcement Learning (RL) and Imitation Learning (IL) face considerable challenges in this post-training phase.

Reinforcement Learning, while capable of producing robust policies through online exploration, is notoriously unstable when fine-tuning VLAs. This instability stems from issues like inaccurate value estimates, especially in complex action spaces, and sparse supervision during intermediate steps. On the other hand, Imitation Learning is easier to train but often suffers from compounding errors, where small mistakes can lead to complete task failure because the robot encounters states not seen in its initial training data.

Addressing these critical limitations, researchers have introduced a novel approach called Human-in-the-loop Online Rejection Sampling (Hi-ORS). This method offers a simple yet highly effective way to achieve both training stability and high robustness in robotic manipulation tasks. Hi-ORS tackles the instability of RL by filtering out negatively rewarded samples during online fine-tuning, thereby stabilizing value estimation. It also employs a reward-weighted supervised training objective, which provides dense, step-by-step guidance to the VLA model.

How Hi-ORS Works: A Closer Look

At its core, Hi-ORS leverages the concept of rejection sampling, a technique widely used in other AI fields like Large Language Models. Instead of relying on potentially inaccurate neural networks to estimate the value of actions, Hi-ORS performs outcome-based filtering. This means it discards robot attempts that result in negative rewards and only retains successful episodes, as judged by a clear reward model. This direct filtering based on task success provides reliable feedback for policy improvement.

Furthermore, Hi-ORS incorporates a reward-weighted supervised learning objective. This objective ensures that the robot receives dense supervision not just for its final actions, but also for the intermediate steps of its inference process. For example, in models that generate actions through iterative denoising, Hi-ORS provides feedback at each denoising step, leading to more efficient and stable learning.

The Human Touch: Guiding Error Recovery

A distinctive feature of Hi-ORS is its seamless integration of human intervention. Recognizing that purely autonomous exploration can be prohibitively expensive and slow for complex tasks, Hi-ORS allows human operators to intervene at any point during a robot’s rollout. These interventions can include teleoperated corrections, targeted resets, or brief corrective segments injected mid-trajectory. Crucially, only human-corrected episodes that achieve positive rewards are retained for training, ensuring that suboptimal human guidance doesn’t contaminate the learning data.

These human corrections serve a dual purpose: they provide efficient exploration guidance, steering the robot towards promising regions of the state space, and, more importantly, they offer explicit demonstrations for error recovery. By showing the robot how to recover from failure modes that would be difficult to discover autonomously, humans help the policy learn robust error-recovery behaviors, significantly boosting its real-world performance.

To maximize data efficiency, Hi-ORS also uses an adaptive interaction frequency. It logs transitions at a higher frequency during human intervention periods to capture fine-grained corrective behaviors, and at a lower frequency during autonomous execution to maintain consistent policy performance.

Also Read:

Real-World Success and Scalability

The effectiveness of Hi-ORS has been rigorously validated across three challenging real-world tasks (Raise-Hand, Pack-Detergent, and Insert-Moisturizer) using two different robotic embodiments (Paxini Tora One and Dobot X-Trainer). In these experiments, Hi-ORS fine-tuned a base VLA policy to master contact-rich manipulation in just 1.5 hours of real-world training. It consistently outperformed traditional RL and IL baselines, demonstrating superior effectiveness and efficiency.

Notably, policies fine-tuned with Hi-ORS exhibited strong test-time scalability. This means the robot could reliably execute complex error-recovery behaviors, using additional retries to recover from intermediate errors rather than repeating failures, leading to higher success rates over multiple attempts. This capability for purposeful recovery is a significant advancement over methods like behavior cloning, which showed limited scaling effects.

The research paper, titled “Human-in-the-loop Online Rejection Sampling for Robotic Manipulation,” authored by Guanxing Lu, Rui Zhao, Haitao Lin, He Zhang, and Yansong Tang, highlights Hi-ORS as a robust and efficient method for fine-tuning VLAs in real-world robotic manipulation tasks. You can read the full paper here.

By combining the stability of rejection sampling with the invaluable guidance of human intervention, Hi-ORS offers a promising path toward more capable and reliable robots in complex, real-world environments.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -