Glossary

Perplexity (Language Model Metric)

A measure of how well a language model predicts a text sample — lower perplexity means the model assigned higher probability to the actual words, indicating better predictive performance.


What it means

Perplexity is an evaluation metric for language models that quantifies how surprised the model is by a held-out text sample. Formally, it is the exponent of the average negative log-likelihood per token:

Perplexity = exp(−(1/N) × Σ log P(token_i))

Intuitively: if a model assigns very high probability to each actual next word in a test text, perplexity is low. If the model is uncertain — spreading probability across many possible continuations — perplexity is high.

Typical ranges:

  • A model with perplexity 1 would perfectly predict every token (impossible in practice)
  • Early GPT-2 achieved ~35 perplexity on Penn Treebank
  • Modern large language models achieve single-digit perplexity on standard benchmarks
  • A uniform distribution over a vocabulary of 50,000 tokens has perplexity of 50,000

Perplexity is commonly used to:

  • Compare models trained on the same data distribution
  • Track training progress
  • Evaluate domain adaptation (a model fine-tuned on medical text should have lower perplexity on medical text)

Why it matters for researchers

Perplexity measures predictive quality, not usefulness. A model can have very low perplexity (excellent at predicting text) while also confidently hallucinating facts. Perplexity is a proxy for how well a model has learned the statistical patterns of text — it does not directly measure accuracy, factuality, or reasoning ability.

Relevance when reading ML papers: When you encounter perplexity scores in language model papers, they are measuring statistical fit on a held-out test set. Lower is better for comparing models trained on the same data with the same vocabulary.

Cross-domain perplexity as a probe: Researchers sometimes measure perplexity on out-of-distribution text to characterize what a model has and hasn’t learned. A model pre-trained only on general web text will have much higher perplexity on specialized scientific notation or code than a model trained with those domains included.

Note: This term refers to the statistical metric. The tool Perplexity.ai is an AI-powered search product that happens to share this name — it is a distinct product, not a direct implementation of this metric.