Pub-AI: AI in Science

JEPA · 2023-01-19

Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas

arXiv:2301.08243PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

I-JEPA learns image representations by predicting the representations of several target blocks of an image from a single context block, in representation space and with no hand-crafted data augmentations.

Why it matters

The image-based starting point of the JEPA (joint-embedding predictive architecture) family: a non-generative alternative to pixel reconstruction that, per the authors, converges faster and learns higher-level semantic features.

Method

  • From one context block, predict the representations of multiple target blocks of the same image.
  • Masking strategy is key: sufficiently large target blocks and an informative, spatially distributed context block.
  • Vision Transformer backbone; no view augmentations during pretraining.

Limitation

Computing targets in representation space adds overhead: about 7% slower time per iteration than pixel-reconstruction methods such as MAE.

Lineage

Abstract (from arXiv)

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.

Related papers