Driving · 2026-09-30
Achieve What You Imagined: Learning to Align Actions with Visual Plans
Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou, Sichao Liu
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors treat a world-action model’s visual predictions as a goal-conditioned visual proposal and use a frozen action-conditioned world model to predict action consequences. They compute dense feedback from cross-model prediction consistency plus terminal goal alignment, then use Flow Policy Optimization to update only the action head. Across four real-world UR5 tasks, they report mean success increasing from 43.4% to 75.1%, without online robot interaction or task-specific reward models.
Why it matters
The paper proposes a way to turn discrepancies between a world-action model’s imagined videos and the consequences of its generated actions into dense rollout-level training feedback. This enables offline post-training that does not require online interaction, extra task-success detectors, or additional task-specific reward-model training, while optimizing a flow-based policy under a pretrained high-order multi-step sampler.
Method
- Use a WAM goal-conditioned video prediction V^C together with generated actions, then treat V^C as a visual goal-conditioned proposal rather than an executable plan.
- Apply a frozen action-conditioned consequence predictor (Ctrl-World) to estimate what the generated actions would actually produce; form feedback from cross-model consistency and terminal goal alignment.
- Optimize only the WAM action head with Flow Policy Optimization (FPO), using the sampler as a black box (native UniPC sampling) and keeping the feedback-generating predictive models frozen.
Limitation
The paper notes that the reward metric is measured in the model space using a classifier applied to Ctrl-World’s predicted terminal states; later reward declines may indicate a mismatch between the optimized policy and the frozen world model’s prediction distribution rather than genuine degradation.
Abstract (from arXiv)
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $\pi_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/
Related papers
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
- EVO-WAM: Evolving World Action Models through Video-Action Verification
- Dynamic Manipulation with World-Action Models via Counterfactual Planning
- WorldLine: Action-Driven Visual Simulation for Robotic Manipulation