JEPA · 2026-09-30
JEPA Learns What the Mask Leaves Unrecoverable
Peng Xie, Amr Alanwar
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors study how JEPA masking geometry determines what can be recovered from context. They model a mask as a linear measurement: atoms whose support lies in the hidden region fall in the null space in a wavelet basis. JEPA must predict only sufficiency, enabling shortcuts with a moving-average target encoder; removing the shortcut depends on coarse-scale content left unrecoverable and on reachable context. They test predictions in 151 pre-training runs.
Why it matters
The paper provides a measurement account of why different mask shapes (blocks, strips, scattered, whole frames) change learned representations in JEPA-style joint-embedding predictive architectures. It links mask geometry to a null-space notion of recoverability (precomputed before training) and isolates two conditions: enough unrecoverable content and enough reachable context. Reported controls include pixel vs latent targets and a frozen target test that modifies the random-vs-block gap.
Method
- Treat masking in JEPA as a linear measurement: in a compactly supported wavelet basis, hidden-region atoms can lie in the measurement null space and leave no trace.
- Analyze JEPA loss as requiring that encoded context suffice for the target; with a moving-average target encoder, shortcuts can become self-consistent, while a frozen target tests this mechanism.
- Score unrecoverable coarse content before training using ν (and ν·n_T) and evaluate mask-geometry predictions across ImageNet-100 and UCF101 with multiple mask families and controls (pixel vs latent, frozen vs moving target, video collator variants).
Limitation
The authors state the account “is established on ViT-S with ImageNet-100 and UCF101” and “fails on spectrograms,” with Appendix L reporting that the ordering does not follow the proposed null-space measure on AudioSet/ESC-50.
Abstract (from arXiv)
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hidden region falls in the measurement's null space and leaves no trace in the data. The JEPA loss asks only that the encoded context suffice for the target, so a target that a low-level prior can recover admits a shortcut, one the moving-average target encoder can make self-consistent. What removes the shortcut is the coarse-scale content the mask leaves unrecoverable, provided enough context stays within reach of each target. We score that content before training and test the account's distinctive predictions in 151 pre-training runs. On ImageNet-100, strip masks match blocks in area and contiguity yet are recoverable, and they land at 40.3% linear top-1, beside random masks at 40.8%, against 64.3% for blocks; within one geometry family, the placements that leave the least unrecoverable content lose 6.5 points to those that leave the most, over five seed pairs; pixel targets span 7 points where latent targets span 25; and against a frozen target the gap between random and block masks, 19 points on the same kind of GPU, closes to 1.5, so the geometry acts through the target the encoder produces for itself. On UCF101 the masking ratio decides which condition, content or reach, binds; removing whole frames, unrecoverable in space but recoverable from neighbouring frames, is worst at both ratios; and on V-JEPA's own masks, batching them intact instead of truncated changes little (36.0% against 35.1%), whereas making 100 target tokens inside the blocks visible lifts them to 48.7% and hiding 100 context tokens outside the blocks does not (33.7%).
Related papers
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Do-JEPA: From Masking to Intervention in Latent World Models
- What masking geometry works best for EEG foundation models?
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models