Driving · 2026-09-30
EVO-WAM: Evolving World Action Models through Video-Action Verification
Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
EVO-WAM adapts world action models to unseen tasks using self-generated video-action trajectories without external action execution. It enables complete autoregressive rollouts via state prediction and anchored multi-frame context, then selects reliable prefixes using a vision-language model for task-completing visuals and an inverse dynamics model to verify video-action consistency. Iteratively training on verified prefixes improves success on unseen RoboTwin 2.0 tasks and real-world long-horizon composites.
Why it matters
The authors address adapting robot policies to unseen tasks without additional expert demonstrations or external execution feedback. They propose a generate–verify–improve loop for world action models that aims to extract supervision from the model’s own generations while filtering out failures where videos do not depict task completion or where video-action pairs are inconsistent. The reported improvements on multiple unseen tasks (simulation and real world) suggest that verification of both task completion and video-action consistency can make self-training effective for world action models.
Method
- Enable autoregressive rollouts by training the WAM to predict end-of-chunk robot state and using an initial visual anchor plus recent generated frames as context for continuation.
- Select reliable self-training data via two-stage video-action verification: a VLM identifies task-completing prefixes, then an inverse dynamics model (IDM) checks video-action consistency by reconstructing actions from generated video and comparing to WAM-generated actions.
- Iteratively generate candidates, verify prefixes, accumulate accepted data across rounds, and supervised fine-tune the WAM on retained verified prefixes plus original base training data.
Limitation
The large early gains suggest useful supervision from initial model generations, while later rounds bring smaller gains and occasional regressions, showing that additional self-training does not always improve performance.
Abstract (from arXiv)
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.
Related papers
- Achieve What You Imagined: Learning to Align Actions with Visual Plans
- Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
- WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation