Pub-AI: AI in Science

Driving · 2026-09-30

Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

Xiangcheng Zhan, Zirui Chen, Yicheng Zhao, Ziteng Gao, Shuo Yang

arXiv:2609.37398PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

DEWO (Direct Experience World-Model Optimization) is a post-deployment learning paradigm for World-Action Models that refines world representations using visual futures collected around interaction turning points. It trains from matched successful and failed continuations for outcome-conditioned visual prediction, and uses a progress-estimating value head to activate classifier-free guidance when progress stalls. Across DexJoCo tasks and real-world dexterous-hand grids, the authors report improved control and predictive losses.

Why it matters

The authors address a deployment gap: post-deployment learning often improves behavior without explicitly improving world predictions, which can be problematic in dexterous manipulation where small errors compound after interaction turning points. DEWO makes world-model adaptation a direct part of control by using observed visual futures (including failures) to condition action generation, and by selectively applying guidance based on estimated progress.

Method

  • Collect experience around interaction turning points by replay exploration, retaining matched successful and failed continuations from comparable contexts.
  • Train outcome-conditioned world-model learning: successful continuations supervise visual and action prediction under a success-conditioned condition, while failed continuations supervise visual prediction under a failure-conditioned condition.
  • At inference, use a value head on video representations to gate classifier-free guidance toward the learned success condition when estimated progress stalls.
Abstract (from arXiv)

World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.

Related papers