Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Control-Geometry Straightening for Sampling-Based Latent Planning

Ziang Fu, Ning Ning

arXiv:2609.35603PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors introduce Control-Geometry Straightening (CGS), a single auxiliary loss for joint-embedding latent world models. CGS uses only local pixel–action transitions to match pairwise cosine similarities among actions to those among corresponding latent differences, shaping “control geometry” for sampling-efficient latent planning. Under linear-dynamics, they analyze how this relates to temporal straightening, finite-budget guarantees for MPPI, and convergence results for gradient descent; experiments show large success-rate gains over LeWM and LeWM+TS.

Why it matters

Accurate transition prediction alone does not ensure that planning objectives are easy to optimize. CGS targets the geometry exposed to action optimization by straightening how action effects accumulate in latent space, aiming to improve nonconvex planning under tight sampling budgets. The paper provides theory connecting learned geometry to planning curvature/conditioning and finite-budget behavior, plus experiments showing success-rate improvements and better sample efficiency.

Method

  • Train a latent world model with a joint prediction loss (LeWM) plus CGS: match pairwise cosine similarities among standardized actions to those among corresponding normalized latent differences (consecutive latent deltas) using only local pixel–action transitions.
  • Use CGS as an auxiliary loss across end-to-end learned (LeWM) or pretrained representations; during planning, optimize an action sequence by minimizing terminal goal-matching cost with the frozen predictor and a sampling-based planner.
  • Provide linear-dynamics theory linking CGS to temporal straightening/control isotropy, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent.
Abstract (from arXiv)

Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to those among corresponding latent differences only using local transitions from pixel-action pairs. The loss can be applied across world-model architectures using end-to-end learned or pretrained representations. Under linear-dynamics, our theoretical analysis connects this objective to temporal straightening and more balanced terminal-cost curvature across the full planning horizon, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent. Across four control environments and multiple planners, CGS improves planning with fewer sampled candidates and refinement steps, achieving success-rate gains up to 20 and 12.6 percentage points over LeWorldModel (LeWM) and its temporal-straightening variant (LeWM+TS), respectively, with sampling-based planners using 128 candidates per update. Probes, comparisons with DINO-WM architecture, and planner-side ablations clarify how latent motion organization, state dependence, and dynamical context shape planning behavior. Straightening control geometry thus makes good action sequences easier to find under limited planning budgets.

Related papers