JEPA · 2024-02-15
Revisiting Feature Prediction for Learning Visual Representations from Video
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
V-JEPA trains video models solely with a feature-prediction objective (predicting masked spatio-temporal regions in representation space), with no pretrained image encoders, text, negatives or reconstruction, on 2 million public videos.
Why it matters
Extends JEPA from images to video and shows frozen features work on both motion-based and appearance-based tasks. The authors report the best performance among methods considered on Something-Something-v2 (+6% accuracy) and a large gain on motion understanding over image models.
Method
- Predicts features of masked video regions from visible context; the target encoder is a moving average of the online one.
- Trained only on 2 million videos from public datasets.
- Evaluated with frozen backbones and attentive probes on image and video tasks.
Lineage
- builds on → Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (Extends predicting in representation space (I-JEPA) from images to video.)
- ← builds on by V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (Scales the V-JEPA mask-denoising objective and adds action-conditioned post-training.)
Abstract (from arXiv)
This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.
Related papers
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- JEPA Learns What the Mask Leaves Unrecoverable
- Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning