Pub-AI: AI in Science

JEPA · 2026-09-30

What masking geometry works best for EEG foundation models?

Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi, Arnaud Delorme, Thomas Moreau, Michael Tangermann

arXiv:2609.33487PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors run a controlled sweep over spatio-temporal EEG masking geometry for 58 masked-prediction foundation models trained with MAE and JEPA, then evaluate all models on 12 OpenEEGBench datasets using linear probing. Both frameworks agree on an optimal mask configuration and shared failure modes; performance is robust outside the best region, while JEPA has an additional failure mode (“bias-inflation collapse”).

Why it matters

Masking determines what an EEG foundation model must predict and from what context, but prior work did not isolate masking geometry from other pipeline changes. The authors provide a unified masking parameterization, evaluate it under two self-supervised frameworks with one shared pipeline, and identify both robust operating regions and JEPA-specific failure behavior that standard detectors miss.

Method

  • Unify EEG spatio-temporal masking strategies with parameters (spatial radius r, temporal length L, target mask ratio ρ*) and sweep a grid of (L,r) configurations (29 per framework) at fixed ρ*=0.55.
  • Train 58 pre-trained EEG foundation models under a shared REVE-Small backbone and dataset/corpus/training recipe, varying only (i) SSL framework (MAE vs JEPA) and (ii) masking geometry.
  • Evaluate all checkpoints on OpenEEGBench with a linear probe on frozen features, using dataset normalization, multi-seed averaging, and statistical testing with two-level bootstrap.
Abstract (from arXiv)

EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.

Related papers