Researchers have developed a system named PRISM that addresses a key challenge in training humanoid robots: the scarcity of diverse real-world data for teaching complex skills. The PRISM framework, detailed in a recent arXiv paper, uses counterfactual video generation to create a vast training dataset from a limited number of real human-object interaction videos. This method allows humanoid robots to learn loco-manipulation skills, such as grasping and transporting objects, without requiring extensive real-world demonstrations or fine-tuning on the physical robot.
The core innovation of PRISM lies in its ability to generate "counterfactual" videos. These are not simply copies of existing footage but rather plausible alternative interactions derived from a few initial real-world examples using video-to-video generation techniques. By creating hundreds of these synthetic yet realistic scenarios, PRISM significantly expands the diversity of training data. This amplified dataset is then used to train a single policy capable of generalizing to unseen objects within specific categories.
Following the video generation phase, PRISM employs a contact-anchored real-to-sim pipeline. This process reconstructs human and object motions from the videos, transforming imperfect visual data into physically plausible trajectories suitable for robot training. The system reconstructs the camera, human motion, and object geometry in a shared world frame. It uses contact points to constrain object pose optimization, which is particularly important when dealing with the ambiguities inherent in monocular video. These reconstructed motions are then retargeted to a robot, preserving interaction phases and aligning robot end-effectors with intended contact points.
The retargeted demonstrations train a privileged teacher policy, which is subsequently distilled into a student policy. This student policy is depth-based and conditioned on onboard depth sensor data and joystick commands, enabling zero-shot deployment onto a real robot. The researchers demonstrated PRISM's effectiveness by deploying a trained policy on a physical humanoid robot. The robot successfully executed pick-up, carry, and drop behaviors for objects including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations, relying solely on its onboard depth observations.
This approach contrasts with traditional methods that often require extensive teleoperation, motion capture, or large-scale simulation setups, which can be costly and time-consuming to scale. While other research has explored sim-to-real transfer for humanoid loco-manipulation using domain randomization and large-scale simulations, PRISM's unique contribution is its method for generating diverse training data through counterfactual video synthesis, thereby overcoming the practical barrier of collecting high-quality, varied real-world interaction videos. The PRISM framework aims to pave the way for more generalist robots capable of performing a wider range of tasks in real-world environments.
