Pub-AI: AI in Science

JEPA · 2026-09-30

Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling

Ch Muhammad Awais, Marco Reggiannini, Davide Moroni

arXiv:2609.31657PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).

Why it matters

The paper targets how foundation-model embeddings behave under severe long-tail class imbalance and evaluates a training-efficient imbalance mitigation approach that avoids end-to-end fine-tuning. It also emphasizes that improvements in aggregate metrics (accuracy, Macro-F1) can coincide with persistent class-wise failures on rare categories, motivating per-class reporting.

Method

  • Benchmark DOFA and SAR-JEPA on imbalanced OpenSARShip under a fixed, training-efficient protocol that keeps the backbone frozen; extract embeddings once and train only a lightweight classifier head.
  • Apply four minority-only embedding-space oversampling methods (SVM-SMOTE, KMeans-SMOTE, SMOTE-ENN, ADASYN) to minority classes until each reaches 3x its original size; leave validation/test embeddings untouched.
  • Train a small 3-layer MLP classifier head (batch norm, ReLU, dropout p=0.1) with cross-entropy and Adam; report results averaged over three random seeds. Code is provided for reproducible multi-seed evaluation.

Limitation

Results are averaged over three random seeds with a fixed dataset split, so the reported variability reflects initialization and sampler randomness but does not capture split sensitivity. Oversampling uses a single minority threshold and a fixed target multiplier for all classes, and feature-space interpolation assumes linear combinations of neighbors remain plausible within the representation space, which may not hold equally for all classes and backbones.

Abstract (from arXiv)

Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with ImageNet-pretrained baselines under a fixed, training-efficient protocol that keeps the backbone frozen. To mitigate imbalance without fine-tuning, we apply four oversampling methods in embedding space exclusively to minority classes and train a lightweight classifier head on the augmented embeddings. Across both foundation models, oversampling improves Macro-F1 and test accuracy relative to their respective baselines, with the largest Macro-F1 gains observed for DOFA using ADASYN (34.39 to 38.56) and for SAR-JEPA using SVM-SMOTE (25.89 to 32.30). We also report class-wise behavior, showing that aggregate improvements can coexist with persistent failures on specific rare classes. Code for embedding extraction and reproducible multi-seed evaluation is provided to support rapid experimentation on free-tier hardware.

Related papers