p-value
The probability of observing results at least as extreme as the data if the null hypothesis were true — widely misunderstood, widely misused, and insufficient on its own for evaluating scientific evidence.
What it means
A p-value is the probability of obtaining test results at least as extreme as the observed results, assuming the null hypothesis is true. A p-value of 0.03 means: if there were truly no effect (null hypothesis), there is a 3% chance of observing data this extreme or more extreme by chance alone.
What p-values do not tell you:
- The probability that the null hypothesis is true
- The probability that your result will replicate
- The size or importance of the effect
- Whether the result is scientifically meaningful
The p < 0.05 threshold for “statistical significance” is a convention, not a law of nature. It was adopted arbitrarily and has caused significant harm through selective reporting, multiple testing problems, and the conflation of statistical and practical significance.
The p-value and the replication crisis
Much of the replication crisis in psychology, medicine, and social science traces to p-value misuse:
- p-hacking: Running multiple analyses and reporting only the one that crosses 0.05
- HARKing (Hypothesizing After Results are Known): Presenting exploratory findings as confirmatory
- Publication bias: Journals preferentially publishing p < 0.05 results, making the literature unrepresentative
AI tools that mine the literature inherit these biases. A meta-analysis trained on published studies with positive results will overestimate true effect sizes.
What to use instead (or alongside)
Effect size — how large is the effect, in practically meaningful units? A p < 0.001 result with a Cohen’s d of 0.05 is statistically significant but practically trivial.
Confidence intervals — the range of effect sizes consistent with the data. More informative than a binary significant/not-significant judgment.
Pre-registration — specifying hypotheses and analysis plans before collecting data eliminates most opportunities for p-hacking.
Bayesian methods — provide posterior probabilities that directly answer “how likely is this effect to be real, given what we know?”