Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

One-Step Next-Latent Prediction Is Not a World Model

Shitong Wang, Zhongang Cai, Yuzhou Hong

arXiv:2609.36227PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors argue that next-latent (one-step) prediction is not, by itself, a world model. They show that one-step regression identifies only a conditional mean (a point), while multi-step open-loop error depends on the residual innovation covariance and grows with horizon after perfect one-step fit. They analyze cases with nonlinear means and non-injective observations, and show that an isotropy penalty depends only on the embedding marginal and has zero transition gradient.

Why it matters

This paper formalizes when one-step next-latent losses do and do not yield the transition kernel needed for rollouts. It provides counterexamples where composing one-step predictions fails (nonlinear conditional mean) or where a memoryless one-step map cannot determine future observations (non-injective observations), and it explains why isotropy regularization cannot directly identify the transition dynamics at fixed encoder.

Method

  • Define risks for one-step regression R1(f)=E||f(zt)-zt+1||^2 and open-loop rollouts RK(f)=E||f^(K)(zt)-zt+K||^2; relate identification of conditional mean vs residual covariance.
  • Analyze linear-Gaussian Markov latents, proving that while one-step fit fixes R1, open-loop squared error grows with horizon K via a trace of pushed-forward innovation covariances.
  • Study nonlinear conditional means, partial observations, and isotropy penalties; prove that isotropy has zero transition gradient with a stop-gradient setup (Theorem 6).
Abstract (from arXiv)

Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.

Related papers