JEPA · 2026-09-30
Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling
Ch Muhammad Awais, Marco Reggiannini, Davide Moroni
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).
Why it matters
The paper targets how foundation-model embeddings behave under severe long-tail class imbalance and evaluates a training-efficient imbalance mitigation approach that avoids end-to-end fine-tuning. It also emphasizes that improvements in aggregate metrics (accuracy, Macro-F1) can coincide with persistent class-wise failures on rare categories, motivating per-class reporting.
Method
- Benchmark DOFA and SAR-JEPA on imbalanced OpenSARShip under a fixed, training-efficient protocol that keeps the backbone frozen; extract embeddings once and train only a lightweight classifier head.
- Apply four minority-only embedding-space oversampling methods (SVM-SMOTE, KMeans-SMOTE, SMOTE-ENN, ADASYN) to minority classes until each reaches 3x its original size; leave validation/test embeddings untouched.
- Train a small 3-layer MLP classifier head (batch norm, ReLU, dropout p=0.1) with cross-entropy and Adam; report results averaged over three random seeds. Code is provided for reproducible multi-seed evaluation.
Limitation
Results are averaged over three random seeds with a fixed dataset split, so the reported variability reflects initialization and sampler randomness but does not capture split sensitivity. Oversampling uses a single minority threshold and a fixed target multiplier for all classes, and feature-space interpolation assumes linear combinations of neighbors remain plausible within the representation space, which may not hold equally for all classes and backbones.
Abstract (from arXiv)
Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with ImageNet-pretrained baselines under a fixed, training-efficient protocol that keeps the backbone frozen. To mitigate imbalance without fine-tuning, we apply four oversampling methods in embedding space exclusively to minority classes and train a lightweight classifier head on the augmented embeddings. Across both foundation models, oversampling improves Macro-F1 and test accuracy relative to their respective baselines, with the largest Macro-F1 gains observed for DOFA using ADASYN (34.39 to 38.56) and for SAR-JEPA using SVM-SMOTE (25.89 to 32.30). We also report class-wise behavior, showing that aggregate improvements can coexist with persistent failures on specific rare classes. Code for embedding extraction and reproducible multi-seed evaluation is provided to support rapid experimentation on free-tier hardware.
Related papers
- $\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
- Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?
- JEPA Learns What the Mask Leaves Unrecoverable