JEPA · 2026-09-30
MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
MA-JEPA introduces a stochastic joint-embedding predictive world model for multi-agent RL with centralized training and decentralized execution. It replaces observation reconstruction with prediction of target observation embeddings, using a categorical latent state and a causal Transformer. A training-only joint predictor conditions on all agents’ local states/actions to predict each agent’s next embedding, which is passed through the same local posterior as in execution for actor-critic learning from latent imagination.
Why it matters
The authors target a core multi-agent issue: local models conditioned on one agent cannot fully predict outcomes depending on other agents’ actions, while centralized latent states can mismatch what is available at decentralized execution. MA-JEPA aims to avoid this by using the same local posterior interface for both real interaction and centralized imagination, while a joint predictor learns interaction effects only during training.
Method
- Local JEPA world model per agent: a categorical stochastic latent state inferred by an observation-conditioned posterior over a causal-Transformer history state.
- Joint training-only predictor: conditions on synchronized local states and joint actions to predict each agent’s next local observation embedding; no centralized actor state is constructed.
- Centralized critic with decentralized execution: critic uses all local states for value learning; neither the joint predictor nor critic runs during execution.
Abstract (from arXiv)
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
Related papers
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control