Confidence Interval
A range of values consistent with the observed data at a given confidence level — more informative than a p-value because it shows both the direction and plausible magnitude of an effect.
What it means
A confidence interval (CI) is a range of parameter values that are statistically compatible with the observed data at a specified confidence level (usually 95%). A 95% CI of [1.2, 3.4] for an odds ratio means the data are consistent with the true odds ratio being anywhere from 1.2 to 3.4.
The correct interpretation — and it is subtle: a 95% CI does not mean “there is a 95% probability the true value falls in this interval.” It means “if we repeated this study many times and calculated this interval each time, 95% of the intervals would contain the true value.” The probability statement is about the procedure, not the specific interval. (The Bayesian analog — a credible interval — does allow the direct probability interpretation.)
Wide vs. narrow intervals:
- A wide CI reflects uncertainty — the data are consistent with a wide range of effect sizes
- A narrow CI reflects precision — the data strongly constrain where the true value lies
- A CI that includes zero (for a difference) or one (for a ratio) is consistent with no effect, regardless of the p-value
CI vs. p-value: why CIs are preferred
A p-value gives a binary signal (significant / not significant). A CI gives continuous information:
- The center of the CI is the best estimate of the effect
- The width of the CI reflects the precision of that estimate
- Whether the CI includes the null value directly answers whether the result is significant at the corresponding alpha level
The American Statistical Association and most major journals now recommend reporting CIs alongside or instead of p-values for precisely this reason.
CIs in AI and machine learning contexts
Prediction intervals vs. confidence intervals. In ML, a prediction interval covers where a new observation is likely to fall (wider, includes both estimation uncertainty and natural variation). A confidence interval covers where the true mean prediction is (narrower). The distinction matters when reporting model performance.
Bootstrap CIs. When the sampling distribution of a statistic is not known analytically, bootstrap resampling generates an empirical CI. This is the standard approach for CI estimation in many ML evaluation settings.
AI extraction of CIs. When using Elicit or AI assistants to extract results from papers, CIs are often extracted less accurately than point estimates. The specific notation (parenthetical, bracketed, reported as ± standard error rather than CI) varies across journals and statistical traditions. Verify extracted CIs against the source.