Pub-AI: AI in Science

Robotics · 2026-09-30

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia

arXiv:2609.38059PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

WorldLine is an action-driven visual simulator for robotic manipulation. It learns transferable robot–object dynamics from 10,000+ hours of action-free robot videos, then grounds heterogeneous actions using 2,000+ hours of action trajectories across 10+ embodiments via an image-space action interface and multi-view/failure-enriched relational training. A robot-focused few-step distillation enables efficient causal rollout for policy evaluation and planning.

Why it matters

The authors target action-faithful, interaction-coherent visual simulation for policy evaluation and embodied planning, addressing two issues: (1) video generation models may not faithfully follow commanded actions and coherent robot–object dynamics, and (2) action-conditioned simulators often require scarce embodiment-specific action data with incompatible control spaces. WorldLine decouples scalable dynamics learning from heterogeneous action grounding and evaluates across action-conditioned generation, policy evaluation, and planning.

Method

  • Three-stage pipeline: Stage I adapts a pretrained TI2V video model to 10,000+ hours of action-free robot videos; Stage II post-trains with geometrically verified action trajectories; Stage III distills into causal few-step rollout.
  • Shared image-space action interface: projects embodiment-specific controls (end-effector pose and gripper state) into each camera view as action maps, avoiding incompatible native control vectors across embodiments.
  • Improves action-conditioned coherence using synchronized multi-view, failure-enriched training, and relational dynamics regularization; then applies robot-focused few-step distillation for efficient causal rollout.

Limitation

In downstream planning, the achievable improvement is also bounded by the diversity of trajectories proposed by the underlying policy and the quality of the multimodal selector.

Abstract (from arXiv)

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{https://zhengsh123.github.io/WorldLine/}{project page}.

Related papers