Pub-AI: AI in Science

JEPA · 2026-09-30

Anisotropic Representations Improve Planning in JEPA World Models

Mingu Kang, Yoori Oh, Sookyung Kim, Joonseok Lee

arXiv:2609.37441PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors show that joint training of JEPA-style latent world models with isotropic Gaussian regularization can learn a representation geometry whose Euclidean latent planning cost ranks feasible outcomes differently from the task cost, even with accurate prediction. They propose AnisoWM with ΛReg: replacing the fixed isotropic target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. Across four visual goal-planning environments, it improves planning success over LeWorldModel in all four.

Why it matters

Latent goal planning often uses Euclidean distance in representation space. This paper argues that the representation regularizer, not just prediction accuracy, determines how terminal errors are weighted during planning, and can create a prediction–planning separation. Learning an anisotropic (diagonal) target reshapes the latent geometry without changing the Euclidean planner, offering a way to align latent planning costs with task outcomes.

Method

  • Introduce AnisoWM with ΛReg: replace SIGReg’s fixed isotropic Gaussian target with a learnable diagonal Gaussian target covariance (Λ = diag(v1,...,vD)) under fixed trace and condition-number anisotropy constraints.
  • Keep the prediction objective, predictor architecture, and Euclidean planner unchanged; use the target only during training (covariance Λ is discarded for planning).
  • Analyze how isotropic SIGReg selects an inverse-covariance geometry in a Gaussian setting, creating finite-horizon action-ranking mismatch; show anisotropic target learning can counteract this induced weighting.
Abstract (from arXiv)

Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with $\Lambda$Reg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/

Related papers