Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Beyond a single latent space: a dual-latent world model for long-horizon planning

Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang

arXiv:2609.37644PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors report Dual-WM, a dual-latent world model that separates low-level execution from high-level long-range planning using distinct state representations/dynamics and learned macro-actions. They introduce LoRe to supervise self-generated recursive rollouts at both levels with horizon-weighted losses. On five goal-conditioned visual control tasks, the authors report improved long-horizon goal success versus task-wise strongest non-actor-guided baselines.

Why it matters

Long-horizon planning with latent world models can fail due to recursive error accumulation and loss of goal discrimination from distance concentration. Dual-WM’s separation of temporal roles (local primitive-action dynamics vs. long-range macro-action planning) targets both issues. LoRe’s weighted supervision aims to improve recursive consistency at each temporal scale, supporting reliable latent planning and goal evaluation beyond short horizons.

Method

  • Dual-WM separates low-level and high-level state representations and dynamics: a low-level model for primitive actions and a high-level model for learned macro-actions, coupled via a learned projection during subgoal refinement.
  • The authors propose LoRe (Long-Horizon Representation Learning with Weighted Rollout), which supervises self-generated recursive predictions at both levels using exponential horizon weights with separate decay rates for the two temporal scales.
  • Planning is coarse-to-fine: high-level route planning generates latent subgoals, the low-level model refines them into primitive actions for precise execution, then switches to direct low-level goal convergence.
Abstract (from arXiv)

Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.

Related papers