Pub-AI: AI in Science

JEPA · 2026-09-30

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

Yutong Feng, Bowen Liao, See Kiong Ng, Yuxuan Liang

arXiv:2609.31680PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.

Why it matters

The paper targets uncertainty about whether JEPA pretraining reliably improves time-series forecasting or depends on the downstream model. By using a consistent JEPA instantiation across a broad backbone set and comparing w/o JEPA vs w/ JEPA with matched training and identical finetuning hyperparameters, the authors show architecture-specific effects that can change which model appears best on a benchmark.

Method

  • Pretrain a JEPA instantiation (context-to-target embedding prediction with LeJEPA + SIGReg) on forecasting datasets, using the same input-label pairs for pretraining and finetuning.
  • For each JEPA-compatible backbone, compare w/o JEPA (random init) vs w/ JEPA (LeJEPA pretrain then finetune) while keeping finetuning stage, loss, and downstream hyperparameters identical and matching total epochs.
  • Exclude backbones that cannot be decomposed unambiguously into encoder/predictor/decoder; skip forcing an arbitrary split when hidden representation boundaries are unclear.

Limitation

The authors report the observed pattern relating JEPA variability to architecture type as an observation rather than an explanation, stating that confirming it would require experiments.

Abstract (from arXiv)

Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on their benefits remains mixed, and most studies test only a single backbone or a narrow set of architectures, leaving unclear whether JEPA pretraining is a reliable improvement or one that depends heavily on the downstream model. We address this gap through a large scale evaluation of one JEPA instantiation across nine backbones and eleven benchmarks spanning temporal and spatio-temporal forecasting, the most extensive cross architecture assessment of JEPA for time series to date. We find that the benefit of this instantiation varies sharply across backbones, producing consistent gains for some architectures and consistent degradation for others, even on the same dataset. This pattern holds across both task families, indicating the variability is a general property of this instantiation rather than a dataset specific artifact worth accounting for when choosing a backbone in practice.

Related papers