Pub-AI: AI in Science

Driving · 2026-09-30

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik, Carmelo Sferrazza, C. Karen Liu, Guanya Shi, Angjoo Kanazawa

arXiv:2609.38172PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

PRISM uses grounded video-to-video (V2V) generation to create hundreds of counterfactual human–object interaction videos from four real seed videos. A contact-anchored real-to-sim pipeline reconstructs and retargets monocular videos into physically plausible robot–object trajectories. A single depth-based humanoid policy is trained and deployed on a real robot zero-shot (no real-world fine-tuning) using onboard depth for pick–carry–drop on diverse objects.

Why it matters

The authors address a bottleneck in visual imitation for humanoid loco-manipulation: collecting large-scale, high-quality interaction videos with clear full-body views and unoccluded contacts. By amplifying a few real videos into diverse counterfactual interactions, then using contact to regularize reconstruction and retargeting, the approach aims to train one policy that generalizes across unseen object instances (and even absent categories) for zero-shot real-world deployment using onboard depth only.

Method

  • Counterfactual video generation: from four real seed videos, PRISM uses grounded V2V (SeedDance 2.0) with category-level prompts to replace the manipulated object while preserving background/lighting/viewpoint, generating 256 counterfactual videos (boxes, bins, barrels, balls).
  • Contact-anchored real-to-sim: extend CRISP-style monocular reconstruction to dynamic objects using SAM3D geometry/motion; detect contact to jointly regularize human and object motion; retarget using object-frame contact anchors to refine noisy monocular reconstructions into physically plausible humanoid–object trajectories.
  • Policy learning and distillation: train a privileged co-tracking teacher in simulation from full-state trajectories, using contact-aware interaction rewards; distill into a depth-based student conditioned on onboard depth and joystick commands for zero-shot sim-to-real deployment.

Limitation

The authors state: “We model dynamic objects as rigid bodies. Thus the policy can struggle with articulated or deformable objects whose contact geometry may change during interaction. For example, foldable chairs may shift through internal joints, causing unstable grasps or loss of control.”

Abstract (from arXiv)

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

Related papers