JEPA · 2026-09-30
Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing
Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman, Le Nguyen, Manuel Lage Ca\~nellas, Sasan Sharifipour, Miguel Bordallo L\'opez
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
Phenomenon-Graph JEPA is a joint-embedding predictive architecture for contactless cardiorespiratory sensing that uses four processed 1D streams, predicts stopped target embeddings along typed edges, and uses within-episode forward prediction without negative pairs or synthetic augmentation in the base setup. On OMuSense-23, the authors report improved label-efficiency area over matched supervised training (3.91 pp) under one configuration, but no advantage of physiological edge typing versus wrong-edge/all-pairs controls.
Why it matters
The authors address scarce labels in contactless cardiorespiratory sensing by pretraining with a non-contrastive predictive objective that avoids negative pairs and augmentation assumptions. Their controlled comparisons (typed vs wrong-edge vs all-pairs graphs, and optional delay coordinates) separate whether gains come from predictive learning itself versus the physiological prior used to organize cross-sensor predictions.
Method
- Joint-embedding predictive objective on four processed 1D streams, using temporal and band-limited spectral branches and predicting stopped target embeddings along typed graph edges.
- Pretraining uses typed cross-edge prediction and within-episode forward prediction, with no negative pairs or synthetic augmentation in the base configuration; predictors and projectors are discarded for fine-tuning.
- Evaluates the physiological typing hypothesis using matched wrong-edge and all-pairs prediction graphs; optional Takens-inspired delay coordinates are studied via validation.
Limitation
The authors state four limitations: (1) OMuSense comparisons use the same ten test participants and C2 was fixed after C1, so the two positive results are observations under two configurations on one cohort rather than independent replications; (2) within-episode sampling uses annotation-derived episode boundaries; (3) Holm families do not correct every exploratory comparison across the complete project; (4) WESAD and OMuSense differ in task, budgets, participants, window lengths and training schedules, so the contrast does not isolate sensor contact as a causal factor.
Abstract (from arXiv)
Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can exploit the unlabeled signals, yet contrastive methods depend on signal transformations and negative pairs whose validity is uncertain for cardiorespiratory data, where time warping changes breathing rate and distant windows can share the same physiological state. We present Phenomenon-Graph JEPA, a joint-embedding predictive architecture that learns from four processed one-dimensional streams without negative pairs or synthetic augmentation in its base configuration. Each stream is encoded by a temporal convolutional branch and a band-limited spectral branch. During pretraining, the model predicts stopped target embeddings along typed edges, which connect streams assigned to the same physiological phenomenon, and forward in time within a state episode. We treat this physiological typing as a testable hypothesis and compare it with wrong-edge and all-pairs prediction graphs. In the OMuSense-23 dataset, pretraining improves label-efficiency area over matched supervised training by 3.91 percentage points (95% interval 2.08 to 5.80, Holm-adjusted p = 0.006), and by 3.74 points under a second configuration evaluated on the same test participants. However, the wrong-edge and all-pairs controls do not establish a benefit from physiological typing. Optional Takens-inspired delay coordinates improve a validation comparison with learned history, whereas two wrist-only WESAD protocols do not establish a pretraining advantage. The study therefore separates the measured benefit of predictive representations from the physiological prior used to organize their training.
Related papers
- Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?
- What masking geometry works best for EEG foundation models?
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- JEPA Learns What the Mask Leaves Unrecoverable
- $\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning