JEPA · 2026-09-30
ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models
Jingnan Pu, Zi-En Fan, Feng Lian
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
ER-JEPA adds an episodic replay pathway to LLM-JEPA. It stores training token pairs in a memory and, at each step, retrieves relevant stored examples to provide additional supervision for both token prediction and representation alignment. The replay path is removed after training, so inference matches LLM-JEPA.
Why it matters
The authors report that LLM-JEPA can converge in representation alignment yet still make persistent prediction errors and later lose previously correct predictions. ER-JEPA’s replay adds supervision from past training pairs to correct errors and improve accuracy while keeping the same inference cost by removing the replay path after training.
Method
- Extend LLM-JEPA with an episodic replay term: the full objective adds a replay loss weighted by β, applied to selected stored training examples.
- Use a fixed-capacity episodic memory; at each step, a replay strategy selects up to R replay entries (content-based, uniform, or hard) and the model recomputes representations to compute the replay JEPA loss.
- Remove the episodic pathway after training so inference uses only the fine-tuned language model (same procedure as LLM-JEPA).
Limitation
The paper states that the relative weighting between different loss terms (e.g., replay weight β) “currently needs to be selected through grid search,” adding computational cost, and that ER-JEPA inherits the need for datasets with paired views (e.g., text and code).
Abstract (from arXiv)
Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.
Related papers
- Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?
- MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents