Overfitting
When an AI model learns the training data too specifically and fails to generalize to new data — a fundamental challenge when training models on small scientific datasets.
What it means
Overfitting occurs when a machine learning model learns the specific patterns of its training data — including noise and idiosyncratic features — so thoroughly that it performs well on training examples but poorly on new, unseen data. The model has memorized rather than generalized.
An overfit model has low training error and high test error — the gap between the two is called the generalization gap. The opposite problem, underfitting, occurs when the model is too simple to capture the patterns in the data and performs poorly on both training and test sets.
Analogy: A student who memorizes specific exam questions rather than understanding the underlying concepts will perform well on those exact questions but fail when faced with novel problems that apply the same principles differently.
Common symptoms in scientific model development:
- A model trained on one institution’s clinical data that fails when applied to another hospital’s records
- A chemical property prediction model that works well on the compounds in its training set but fails on structurally novel compounds
- A classification model that achieves 95% accuracy in cross-validation but degrades to 70% in prospective testing
Why it matters for researchers
Small scientific datasets are particularly vulnerable. Many research domains involve small, expensive-to-collect datasets — molecular property measurements, clinical trial results, electron microscopy images. These are precisely the settings where overfitting is most dangerous.
Mitigation strategies:
- Cross-validation — splitting data into multiple training/test folds to estimate generalization performance more reliably than a single split
- Regularization — techniques (dropout, weight decay, data augmentation) that penalize model complexity during training
- Held-out test set — keeping a completely separate test set that is never used during model development or hyperparameter tuning; if you tune on the test set, you are overfitting to it
The leakage problem: In scientific ML papers, data leakage — where test-set information inadvertently enters the training process — is a common source of overfit-inflated results. When evaluating claims in the literature, look carefully at how training and test sets were constructed (are they temporally split? Are similar molecules/patients split correctly?).