Pub-AI: AI in Science

Robotics · 2026-09-30

One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

Bang Du, Yichen Xie, Shuqi Zhao, Yuxin Chen, Menglin Wu, Masayoshi Tomizuka

arXiv:2609.36413PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors propose RoboActualizer to turn a pretrained video world model into a robot policy by “actualizing” a single task-conditioned future. A frozen V-JEPA 2.1 encoder provides a prior over plausible future latents; a tiny flow-matching head with two DiT experts learns task-conditioned selection and action realization. The approach is trained on cached latents and deployed with low latency for real-time control.

Why it matters

The work targets the costly step of extracting robot actions from video world models by avoiding fine-tuning the heavy backbone. By freezing the world model and learning only a lightweight actualizer for task-conditioned selection and action readout, the authors report strong simulation and real-world results with far fewer trainable parameters and low inference latency suitable for real-time control.

Method

  • Freeze a pretrained latent video world model (V-JEPA 2.1) as a prior over diverse future latents; train only a small flow-matching actualizer head on cached latents.
  • Model actualization as task-conditioned selection (latent expert predicts future latents conditioned on instruction) plus realization (action expert predicts continuous actions to realize the selected future).
  • Implement the actualizer as two lightweight Diffusion Transformers (DiT) experts in a Mixture-of-Transformers (MoT) flow-matching setup; cache latents offline to enable single-GPU training.
Abstract (from arXiv)

A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.

Related papers