Pub-AI: AI in Science

JEPA · 2026-09-30

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao

arXiv:2609.34375PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.

Why it matters

The paper targets a capacity-allocation issue in reward-free, compact JEPA world models: single-latent representations can entangle action-relevant state with predictable but control-irrelevant context, harming planning. By structurally routing dynamics into a compact predictive latent and using residual context only for reconstruction, the approach aims to improve latent-space planning efficiency and control-relevant representation quality without requiring reward supervision.

Method

  • Functional routing: split representations into a compact predictive latent z_t for dynamics/planning and a learned-query residual-context u_t used only for cross-attention reconstruction; dynamics rollout uses z_t only.
  • Train with SIGReg on the predictive latent (anti-collapse), a temporal VICReg-style regularization on u_t (trajectory stability), and patch reconstruction loss via a cross-attention decoder.
  • Test-time planning uses MPC with CEM by rolling out the action-conditioned predictor autoregressively in z-space and scoring terminal predicted latent distance to the encoded goal.

Limitation

The approach assumes nuisance context is approximately stable over the training horizon; rapidly changing but uncontrollable or irrelevant factors may be routed imperfectly. The decoder and context regularizer add training-time computation and introduce sensitivity to bottleneck capacity and loss weights, and real-world data is evaluated via offline action recovery rather than closed-loop robot execution.

Abstract (from arXiv)

Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.

Related papers