Video World Models · 2026-09-30
World2Motion: Turning Video World Models into 3D Human Motion Generators
Tu Fangyuan, Xiangyue Zhang, Yiyi Cai, Yichen Peng, Kunhang Li, Bo Zheng, Zhixiang Wang, Kaipeng Zhang, Erwin Wu, Haoran Xie, Haiyang Liu
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
World2Motion adapts Cosmos 3 into a single-stage generator that produces scene-aware 3D human motion and corresponding video from a single image and a text prompt. The authors build a hybrid training dataset (synthetic video–motion pairs plus real videos paired with estimated 3D motion) and introduce a shift-decoupled noise schedule to reduce motion jitter. On a multi-source interaction benchmark, it improves motion–text alignment and scene interaction and matches two-stage interaction success with ~3.3× faster inference.
Why it matters
The authors aim to avoid the costly two-stage pipeline of generating video then recovering motion. By converting a pretrained video world model into a single-stage 3D motion generator, World2Motion targets better scene-aware full-body motion with lower inference cost and reduced temporal instability (motion jitter). The approach includes training-data construction for scarce paired supervision and a modality-specific noise schedule to improve motion stability in joint video–motion generation.
Method
- Adapts Cosmos 3’s MoT backbone for single-stage joint generation of scene-aware 3D human motion and corresponding video from a single image and a text prompt, using CameraHMR for the initial pose and avoiding a separate motion-recovery stage.
- Constructs HOI-mix: combines synthetic video–motion pairs rendered from interaction motion sequences (with known 3D supervision) and real videos paired with estimated 3D motion using CameraHMR.
- Introduces a shift-decoupled noise schedule that assigns different noise shifts to video and motion while sharing denoising progress, and trains/inferences with matching video–motion noise relationships to reduce motion jitter.
Limitation
World2Motion mainly covers single-person motion and simple interactions, leaving more complex scenarios largely unexplored; it is still too slow for real-time use; and accurate root placement and contact remain challenging.
Abstract (from arXiv)
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference.
Related papers
- MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
- PhysWAM: Physically Consistent World Action Model for Autonomous Driving
- HelixWorld: A Real-time Interactive Audio-Visual World Model
- Honeycomb: Constant-Size Scene Memory Representation for Video World Models
- Genie: Generative Interactive Environments