Pub-AI: AI in Science

JEPA · 2026-09-30

Adaptive Latent Capacity for World Models

Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer, Hai Victor Habi

arXiv:2609.32921PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Adaptive LeWorldModel (ALeWM) is a JEPA-style world model that learns to concentrate predictive information into compact prefixes of a wide latent embedding. It uses a sequence-conditioned capacity network to sample a prefix length during training, then regularizes masked embeddings with MixSIGReg against a Gaussian-active/zero-suffix mixture target. The learned ordering supports recursive planning with lower average planning capacity than fixed-width LeWM while achieving higher mean success rates.

Why it matters

The authors address a representational-ordering problem for world models: how to keep a wide latent capacity while making early coordinates the ones most useful for prediction and planning under smaller latent-state budgets. Their method aims to allow a single model to support multiple prefix capacities (without retraining per width) and to enable planning with an episode-fixed selected prefix, improving control success while reducing average planning capacity.

Method

  • ALeWM learns a sequence-conditioned distribution qψ(k|w) over admissible prefix lengths k and samples a prefix during training; the predictor then estimates the full next embedding from the masked prefix embedding.
  • MixSIGReg regularizes masked embeddings against a prior-weighted mixture distribution with Gaussian active prefixes and zeros in the remaining coordinates, encouraging early coordinates to retain predictive information.
  • At deployment, the capacity network selects a modal capacity from initial and goal observations and MPC planning uses that fixed active prefix for recursive prediction and goal comparison.
Abstract (from arXiv)

We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.

Related papers