Pub-AI: AI in Science

Driving · 2026-09-30

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

arXiv:2609.38163PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

ReWAM is a representation-centric world-action model that uses frozen pre-trained DINO features and builds compact temporal world states via a Temporal Representation Bottleneck (TRB). Action-Grounded Representation Shaping (AGRS) routes only action-loss gradients to the TRB, letting the policy shape what the representation encodes. The authors report 93.6% success on RoboTwin 2.0 without generative video pre-training; on RoboDojo they report average score 9.42 and success rate 5.82% without embodied pre-training and 12.29 and 8.28% with ~600 hours.

Why it matters

The authors argue that the representation space is an interface between control and prediction, and that neither reconstruction fidelity nor uncalibrated pre-trained perceptual features by themselves guarantee effective policy learning. ReWAM proposes to explicitly organize and adapt representations for dynamics modeling and control by (1) calibrating DINO features for diffusion-based prediction, (2) temporally bottlenecking them into compact world states, and (3) shaping the representation using action-loss gradients while keeping world prediction loss from updating the representation via routing constraints.

Method

  • Uses frozen DINOv3-L features with Feature Calibration: multi-layer aggregation, normalization, and a representation-aware noise schedule for diffusion-based dynamics modeling.
  • Introduces a Temporal Representation Bottleneck (TRB) that converts calibrated frame-wise DINO features into compact 128-dim tokens via 2×2×2 spatiotemporal patchification, attention, projection, and running normalization.
  • Trains with Action-Grounded Representation Shaping (AGRS): only action-loss gradients through the current-state reach the TRB; world-loss gradients stop at current and future representations; world and action branches are trained with flow matching and an MoT diffusion-transformer architecture.
Abstract (from arXiv)

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

Related papers