Pub-AI: AI in Science

JEPA · 2026-09-30

Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control

Antonio Pariente, Ignacio Boero, Nikolai Matni, Alejandro Ribeiro

arXiv:2609.36305PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors propose a JEPA-style world model where latent dynamics are restricted to a bilinear (then normalized) parameterization, enabling efficient gradient-based planning and structurally enforcing action recoverability to prevent representation collapse. They report that, on standard 2D/3D control tasks, bilinear-parameterized representations match or improve success while reducing planning time by nearly three orders of magnitude.

Why it matters

This paper targets two issues in JEPA-style world models: (1) the flexibility of latent dynamics can be counterproductive for planning/control, and (2) representation collapse. By constraining latent dynamics to a structured bilinear form and enforcing action recoverability via the model structure, the authors claim both stable representations and substantially faster latent planning (via differentiable sensitivities), including longer-horizon and real-time control regimes.

Method

  • Train a JEPA-style world model entirely in latent space without observation reconstruction, but constrain latent dynamics to a bilinear parameterization followed by Cholesky–QR normalization.
  • Use the structured, differentiable dynamics (state/action sensitivities through the normalization) to replace sampling-based planners like CEM with a gradient-based Gauss–Newton planner in latent space.
  • Evaluate longer-horizon planning and moving-goal (continuous replanning) settings, using real-time control protocols where sampling-based methods are too slow.
Abstract (from arXiv)

World models jointly learn latent representations and dynamics that predict how high-dimensional observations evolve under actions. In this work, we propose a JEPA-style world model in which, rather than learning arbitrary latent dynamics, we restrict them to follow a bilinear parameterization. This structure enables efficient planning and control while shifting the modeling burden onto the encoder, encouraging richer representations that expose the controllable geometry of the system. In particular, this structured parameterization allows us to structurally enforce action recoverability, thereby preventing representation collapse by construction. Although prescribing a bilinear parametrization may appear restrictive, we show that a broad class of nonlinear dynamical systems admits a transformation under which the dynamics become bilinear. Empirically, we show across standard 2D and 3D control tasks that representations with bilinear-parameterized dynamics can be learned directly from high-dimensional observations, reducing planning time by nearly three orders of magnitude while retaining or even improving control accuracy. We also propose more demanding regimes of longer-horizon planning and real-time control, and demonstrate that our method succeeds in both, moving JEPA-style world models beyond short-horizon offline planning.

Related papers