Video World Models · 2026-09-30
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
Tamim Zoabi, Ameen Ali, Lior Wolf
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
H-JEPA is an action-conditioned, reconstruction-free world model that separates a wide perceptual code from a fixed orthonormal control state slice. The control state inherits the code’s covariance, then evolves with phase-conditioned dissipative port-Hamiltonian dynamics. Port-inverse consistency (PIC) ties the action readout to the input port transpose, shown to equal a rollout error projection reweighting prediction error. The authors report matching or exceeding reconstruction-free baselines on four pixel-based control benchmarks in ≤10 epochs.
Why it matters
Planning from pixels requires latent dynamics whose state is organized by how actions move the system. The authors argue that prior-based representation stability is not enough for control and introduce a design that fixes the control state geometry via covariance inheritance, while using a port-Hamiltonian predictor with PIC to reweight rollout error along action-relevant directions. This provides an explicit, structured way to align prediction and planning without pixel reconstruction.
Method
- Split perception and control: encode each frame to a wide code h, then use a fixed orthonormal projection U to define control state s = U^T h; the control state inherits code covariance.
- Predict with phase-conditioned dissipative port-Hamiltonian dynamics where the input port has orthonormal columns (Stiefel constraint).
- Add PIC: read the executed action back through the transpose of the same port and show it equals the rollout error projected onto port directions (a parameter-free reweighting).
Abstract (from arXiv)
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most $10$ training epochs, and its largest gain is on OGB-Cube ($91.9$ against $79.3$ percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.
Related papers
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning