JEPA · 2026-09-30
LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.
Why it matters
The paper targets a capacity-allocation issue in reward-free, compact JEPA world models: single-latent representations can entangle action-relevant state with predictable but control-irrelevant context, harming planning. By structurally routing dynamics into a compact predictive latent and using residual context only for reconstruction, the approach aims to improve latent-space planning efficiency and control-relevant representation quality without requiring reward supervision.
Method
- Functional routing: split representations into a compact predictive latent z_t for dynamics/planning and a learned-query residual-context u_t used only for cross-attention reconstruction; dynamics rollout uses z_t only.
- Train with SIGReg on the predictive latent (anti-collapse), a temporal VICReg-style regularization on u_t (trajectory stability), and patch reconstruction loss via a cross-attention decoder.
- Test-time planning uses MPC with CEM by rolling out the action-conditioned predictor autoregressively in z-space and scoring terminal predicted latent distance to the encoded goal.
Limitation
The approach assumes nuisance context is approximately stable over the training horizon; rapidly changing but uncontrollable or irrelevant factors may be routed imperfectly. The decoder and context regularizer add training-time computation and introduce sensitivity to bottleneck capacity and loss weights, and real-world data is evaluated via offline action recovery rather than closed-loop robot execution.
Abstract (from arXiv)
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.
Related papers
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning