TLDR: KUNGFUBOT2 introduces VMS, a unified whole-body controller for humanoid robots that allows them to learn and perform a wide range of complex and dynamic motion skills with high stability and generalization. It achieves this through a hybrid tracking objective, an Orthogonal Mixture-of-Experts architecture for skill specialization, and a segment-level tracking reward for robustness, validated in both simulations and real-world scenarios on the Unitree G1 robot.
Humanoid robots have long captured our imagination, promising a future where machines can mimic human behaviors, from simple walking to intricate martial arts. However, teaching these robots a broad range of dynamic and stable movements within a single control system has been a significant challenge. Researchers have now introduced a groundbreaking framework called VMS (Versatile Motion Skills) as part of the KUNGFUBOT2 project, designed to empower humanoid robots with an unprecedented ability to learn and execute diverse whole-body movements.
The core idea behind VMS is to create a single, unified controller that can master a vast repertoire of motion skills while maintaining stability over extended periods. This is particularly difficult because different movements require distinct control strategies, and ensuring smooth transitions and long-term balance is crucial.
How VMS Achieves Versatile Control
VMS integrates several innovative components to overcome these challenges:
- Hybrid Tracking Objective: This objective balances two critical aspects of motion. It ensures that the robot’s movements closely match the local style and fidelity of the reference motion (like how an arm moves during a throw) while also maintaining global consistency of the overall trajectory (ensuring the robot doesn’t drift off course). This prevents the common issues of either losing local detail or drifting globally.
- Orthogonal Mixture-of-Experts (OMoE) Architecture: Imagine a team of specialized experts, each trained for a different type of movement. The OMoE architecture works similarly, using multiple ‘expert’ networks whose outputs are kept distinct or ‘orthogonal’. A ‘router’ network then intelligently selects and combines these experts based on the current motion, allowing the system to specialize in different skills (like walking, kicking, or dancing) without them interfering with each other. This significantly enhances the robot’s ability to learn and generalize across a wide array of movements.
- Segment-level Tracking Reward: Traditional methods often penalize robots for even slight, momentary deviations from a reference motion, which can lead to instability, especially in dynamic or long sequences. VMS introduces a more forgiving ‘segment-level’ reward. Instead of demanding perfect alignment at every single moment, it rewards the robot for aligning with the most feasible reference state within a short future window. This makes the robot more robust to temporary errors and allows for smoother, more stable execution of long and complex motions.
Also Read:
- Advancing Humanoid Locomotion with Mamba-Powered Reinforcement Learning
- Bridging Human and Humanoid Movement: A New Approach to Motion Retargeting
Real-World Demonstrations and Future Potential
The VMS framework has been rigorously tested in both simulations and real-world experiments using the Unitree G1 humanoid robot. The results are impressive, showcasing the robot’s ability to accurately imitate dynamic skills and maintain stable performance for minute-long sequences. The robot has demonstrated a wide range of capabilities, including various locomotion styles (walking, running), athletic movements (ball throwing, racket swinging), expressive dances, diverse kicking techniques, and even complex sequences from Kung Fu and other martial arts.
Beyond these core demonstrations, VMS also shows strong generalization capabilities. It can be used for text-to-motion generation, where the robot follows motion instructions generated from language descriptions. Furthermore, it adapts well to challenging, out-of-distribution movements like lying down and getting up, or highly acrobatic spinning kip-ups, with minimal fine-tuning.
While VMS represents a significant leap forward in humanoid whole-body control, the researchers acknowledge its current limitations. It currently lacks visual perception, which limits its understanding of complex environments, and it relies heavily on large-scale motion capture datasets, which might not always cover all skills evenly. Addressing these areas will be key for future advancements.
This work, detailed in the paper KUNGFUBOT2: Learning Versatile Motion Skills for Humanoid Whole-Body Control, lays a strong foundation for developing general-purpose humanoid robots capable of performing a vast array of human-like tasks with remarkable agility and stability.


