Pub-AI: AI in Science

JEPA · 2024-02-15

Revisiting Feature Prediction for Learning Visual Representations from Video

Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas

arXiv:2404.08471PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

V-JEPA trains video models solely with a feature-prediction objective (predicting masked spatio-temporal regions in representation space), with no pretrained image encoders, text, negatives or reconstruction, on 2 million public videos.

Why it matters

Extends JEPA from images to video and shows frozen features work on both motion-based and appearance-based tasks. The authors report the best performance among methods considered on Something-Something-v2 (+6% accuracy) and a large gain on motion understanding over image models.

Method

  • Predicts features of masked video regions from visible context; the target encoder is a moving average of the online one.
  • Trained only on 2 million videos from public datasets.
  • Evaluated with frozen backbones and attentive probes on image and video tasks.

Lineage

Abstract (from arXiv)

This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.

Related papers