Latent Dynamics
Dreamer/RSSM-style latent dynamics and model-based RL. (34 papers)
2026-09-30 · Hitoshi Iyatomi · arXiv:2609.32268
The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.
Latent Dynamicsauto-summary
2026-09-30 · Linhao Wang, Yiyan Fan, Dongjin Huang · arXiv:2609.32322
The authors show that two world models with similar total prediction error can yield very different planning outcomes when their errors occur on different state dimensions. They introduce Decision-Relevant Prediction Error (DRPE), measuring multi-step error only on decision-relevant state dimensions, and an iso-error protocol that varies error allocation while keeping total error fixed.
Latent Dynamicsauto-summary
2026-09-30 · Enzo Nicol\'as Spotorno, Josafat Leal Filho, Ant\^onio Augusto Fr\"ohlich · arXiv:2609.32512
The authors present a measurement protocol for action-conditioned latent vehicle predictors with a physical readout. It separately tests retention of physical quantities, organization in latent space, one-step forecasting, and local response to small command perturbations via three matched response paths. In a case study on IPG CarMaker data, representations retain planar outputs, improved future-command inputs help 1s forecasts, but local command-response can diverge in latent space and cause regret in nearby-command ranking.
Latent Dynamicsauto-summary
2026-09-30 · Dai Shi, Andi Han, Feng Chen et al. · arXiv:2609.32966
The authors show that RL’s representation-fitting loop can sustain a lower-return policy: in a self-confirming superposition trap, globally optimal codes under the current policy share overlapping directions for features that rarely co-activate, yet interfere after an alternative action. Interventions that preserve access to neglected states or protect replay weights reduce interference and can improve control, including in DreamerV3–Crafter.
Latent Dynamicsauto-summary
2026-09-30 · Yangyuan Li, Weichao Li, Shaowu Pan · arXiv:2609.33205
SMORE is a mesh-agnostic reduced-order model for time-dependent PDEs. It uses an INR autodecoder to map sparse spatial measurements to a latent state, then evolves the latent state with structured latent dynamics (linear or linear-quadratic). A Lyapunov-guided stability regularization is added to promote stable long-horizon rollouts, and the authors provide theoretical guarantees under structural assumptions.
Latent Dynamicsauto-summary
2026-09-30 · Meng Zhu, Airui Zhang · arXiv:2609.33347
MultiEcho treats frozen world models as experimental systems by estimating response laws from controlled counterfactual interventions. Using “three-reference” forward fits and reverse readouts, the authors predict complete intervention responses and recover intervention parameters across nine simulated physical systems and seven frozen model configurations, then test applicability and physical correspondence via temporal/visual/material intervention variants.
Latent Dynamicsauto-summary
2026-09-30 · Alexis-Raja Brachet, Guillaume Clavier--Fr\'emond, Abdelhakim Ziani et al. · arXiv:2609.33566
The authors introduce Observer State-Space Models (OSSMs) for time-series forecasting. OSSMs treat the observed input sequence as measurements correcting an internal latent state via an observer, while the latent transition is autonomous and shared across context and forecasting intervals. They claim OSSM improves forecasting versus corresponding SSMs under the same parameter count and training setup.
Latent Dynamicsauto-summary
2026-09-30 · Boyuan Zhang, Yingjun Du, Xiantong Zhen et al. · arXiv:2609.33595
The authors introduce SALT, an action-conditioned joint-embedding latent dynamics model with state-affine transitions and recursive rollout training. They argue one-step prediction error can fail to predict planning quality because planning composes transitions recursively, transforming introduced errors. Across four visual planning environments, SALT has higher one-step prediction error than LeWM but higher closed-loop success.
Latent Dynamicscodeauto-summary
2026-09-30 · Jifan Li, Ning Ning · arXiv:2609.33844
ViBR-WM is a visual world model that forecasts in a DINOv2-based latent representation and uses Bayesian regression to predict joint visual–physical states recursively and physical targets directly. It combines trend, seasonal, and cycle modules with interpretable regression and Bayesian variable selection, averaging over predictor subsets via posterior probabilities and modeling parameter uncertainty and future disturbances.
Latent Dynamicsauto-summary
2026-09-30 · Alexander Detkov, Matt Thomson · arXiv:2609.34058
The authors study whether models learn global constraints and propagate their consequences from local transition data. In controlled monoid worlds, next-state training fits paths but fails to propagate non-trivial constraints, while compositional training (hiding intermediates) achieves 96% average accuracy on inverse/commutativity/composition. Generalization declines sharply with proof depth; longer compositional path length improves deeper propagation and improves related results in vision and language.
Latent Dynamicsauto-summary
2026-09-30 · Ji Dai, Quan Fang, Junyu Gao et al. · arXiv:2609.34604
SPRII (Shaping Persistent Representations from Independent Interactions) is a relation-supervised training principle for world models that learns persistent context from related independent interactions without numerical property labels. It uses two components: Align (related contexts agree) and Cross (use one interaction’s context to predict another’s future), while keeping the learner’s native objective. The paper analyzes Formation, Use, and Value links from accessible persistent information to task error reduction.
Latent Dynamicsauto-summary
2026-09-30 · Haowei Xu, Wanyi Fu, Hongbin Han et al. · arXiv:2609.35338
NeuronDiscover is an Agent-in-Twin framework for mechanistic discovery in neuronal microenvironments under “twin confounding” (mechanism change vs computational twin error). A shared mechanism-grounded World Action Model (WAM) drives prediction, intervention proposals, and observation design; independently adjudicated outcomes revise a scoped Mechanism–Intervention–Observation–Outcome (MIOY) graph, compiling supported relations into executable programs with discrepancy-adjusted acceptance bounds. Evaluated in simulated transport worlds and donor-disjoint current-clamp recordings.
Latent Dynamicsauto-summary
2026-09-30 · Ziang Fu, Ning Ning · arXiv:2609.35603
The authors introduce Control-Geometry Straightening (CGS), a single auxiliary loss for joint-embedding latent world models. CGS uses only local pixel–action transitions to match pairwise cosine similarities among actions to those among corresponding latent differences, shaping “control geometry” for sampling-efficient latent planning. Under linear-dynamics, they analyze how this relates to temporal straightening, finite-budget guarantees for MPPI, and convergence results for gradient descent; experiments show large success-rate gains over LeWM and LeWM+TS.
Latent Dynamicsauto-summary
2026-09-30 · Haojian Huang, Zexi Li, Junhao Guo et al. · arXiv:2609.36012
The authors review in-context learning (ICL) for robots, where demonstrations and interaction change deployed behavior without updating neural parameters during deployment. They organize prior work into four interfaces connecting context to execution: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution, analyzing how training, correspondence, and memory determine whether taught requirements transfer under changing objects, environments, and execution conditions.
Latent Dynamicsauto-summary
2026-09-30 · Shitong Wang, Zhongang Cai, Yuzhou Hong · arXiv:2609.36227
The authors argue that next-latent (one-step) prediction is not, by itself, a world model. They show that one-step regression identifies only a conditional mean (a point), while multi-step open-loop error depends on the residual innovation covariance and grows with horizon after perfect one-step fit. They analyze cases with nonlinear means and non-injective observations, and show that an isotropy penalty depends only on the embedding marginal and has zero transition gradient.
Latent Dynamicsauto-summary
2026-09-30 · Ke Fang, Yupu Yao, Lu Cheng · arXiv:2609.36333
ATLAS addresses a gap in latent world-model planning: regularizing only the latent marginal (anti-collapse) does not guarantee preservation of the relational geometry needed for goal-conditioned action selection. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and calibrates the planning latent’s marginal using Wasserstein embedding matching (WEMReg). Instantiated in LeWM, it improves mean goal-reaching success, especially on higher-novelty TwoRoom episodes.
Latent Dynamicsauto-summary
2026-09-30 · Iishaan Inabathini, Margaret M. Henderson · arXiv:2609.36366
The authors extend cross-attention fMRI encoding models from static images to naturalistic video. They use per-parcel joint spatiotemporal cross-attention (joint over space and time) over V-JEPA-2 video tokens, fitting to BOLD Responses from short video clips. Joint routing yields higher held-out prediction accuracy across higher visual regions and produces interpretable attention maps that track moving objects and differ across category-selective networks.
Latent Dynamicsauto-summary
2026-09-30 · Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen et al. · arXiv:2609.36845
DSWM is a decomposed spatio-temporal world model for demand-driven UAV base station repositioning. It uses a recurrent state-space model trained with an EMA-based latent predictive objective plus variance regularization, a differentiable service simulator head (association, probabilistic LoS, Shannon rate, capped fulfillment), and observation-anchored CEM planning that executes only the first action of the best imagined rollout.
Latent Dynamicsauto-summary
2026-09-30 · Ziqi Liu, Songhan Yang, Linfan Zhou et al. · arXiv:2609.36985
The authors propose Abductive World Modeling (AWM), which learns structured causal representations by “predict forward, then abduce backward.” Using the Hierarchical Abductive State Pyramid (HASP), AWM infers a hierarchical latent state with Entity, Dynamic, and Relation components by jointly reasoning over the current observation and its predicted future. Experiments on physical prediction, causal reasoning, and action understanding show consistent gains over a V-JEPA 2 backbone-only baseline.
Latent Dynamicscodeauto-summary
2026-09-30 · Ziqi Wen, Ting Xu, Lianyu Wang et al. · arXiv:2609.37156
LucidWM is a world model that learns “doubt” about imagined transitions from experience using Subjective Logic, then turns doubt into “trust” that compounds along imagination. Trust reweights λ-returns for learning and guides action choice. It needs no extra parameters or forward passes for uncertainty estimation, and the authors report earlier alarms for drift and fewer steps to reach a goal in a navigation study.
Latent Dynamicsauto-summary