Skip to content
All lessons
Cross-cutting · Biostatistics

Confidence Intervals & Statistical Power

A Step 1-focused lesson on interpreting confidence intervals (crossing 0 for difference measures, 1 for ratios) and recognizing statistical power, Type I/II error, and the underpowered \"negative\" study.

11 min readHigh yield

Why This Matters on Step 1

Confidence intervals (CIs) and statistical power are the interpretation tools Step 1 loves, because they turn a bare p-value into a story. A confidence interval is the range of values likely to contain the true population parameter. A 95% CI means that if the study were repeated many times, 95% of the computed intervals would capture the truth — it is NOT "a 95% chance the true value lies in this one interval." Power is the mirror image: the probability that a study can actually detect an effect that is truly there. The boards almost never make you calculate. They make you interpret: does the interval cross the line of no effect, and was a "negative" trial simply too small to find a real difference?

Confidence Intervals — Core Facts
  • 95% CI = point estimate ± 1.96 × SEM, where SEM = SD/√n (99% CI uses 2.58; 90% uses 1.645)
  • Higher confidence = WIDER interval (99% CI is wider than 95%); a wider CI is less precise
  • A CI narrows with ↑ sample size, ↓ standard deviation, or ↓ confidence level
  • Difference measures (mean difference, risk/rate difference): significant if the CI does NOT include 0
  • Ratio measures (RR, OR, HR): significant if the CI does NOT include 1
  • Two groups whose 95% CIs do not overlap differ significantly (p < 0.05); overlapping CIs are inconclusive
  • A CI crossing the null carries the same verdict as p > 0.05 — but also shows effect size and precision
Bar chart of many 95% confidence intervals from repeated samples, with the intervals that fail to capture the true parameter value highlighted
The frequentist meaning of a 95% CI: across many repeated samples, about 95% of the computed intervals capture the true value — it is not a 95% probability that this one interval does. · Wikimedia Commons — Leonhard Bamberg — CC BY-SA 4.0, via Wikimedia Commons
Vignette — Reading a Confidence Interval

Vignette: An RCT of a new antihypertensive vs placebo reports a relative risk of stroke of 0.78 (95% CI 0.62–0.94).

Interpretation: This is a ratio measure, and the CI excludes 1, so the reduction is statistically significant (p < 0.05) — the drug lowers stroke risk.

Contrast: If instead RR were 0.85 (95% CI 0.70–1.05), the interval crosses 1not significant; you cannot conclude benefit despite the point estimate < 1.

Next step in reasoning: Confirm the significant effect is also clinically meaningful (magnitude, NNT), and check that the interval is narrow (precise) enough to trust rather than merely non-null.

Errors & Power at a Glance

ConceptMeaningTypical value
Type I error (α)False positive — reject a TRUE null (see a difference that isn't there)0.05
Type II error (β)False negative — fail to reject a FALSE null (miss a real difference)0.20
Power (1 − β)Correctly detect a true effect that exists≥ 0.80 desired
p-valueChance of data this extreme if the null were truereject null if < α
Statistical Power — Core Facts
  • Power = 1 − β = probability of correctly rejecting a false null (detecting a true effect)
  • Increase power by: ↑ sample size (most controllable), ↑ effect size, ↑ α, ↓ variability (SD), one-tailed test
  • Sample size is the lever investigators actually control — the usual board answer to "how do you boost power?"
  • An underpowered study (small n) risks a Type II (β) error: a false "no difference"
  • For a fixed n, α and β trade off: lowering α (fewer false positives) raises β (more false negatives)
  • "Absence of evidence is not evidence of absence" — a non-significant result ≠ proof of no effect
Overlapping null and alternative hypothesis distributions for a two-sided test, illustrating alpha, beta, and statistical power
Power (1 − β) is the area of the alternative distribution beyond the critical value; shrinking β or separating the curves (larger effect or larger n) raises power. · Wikimedia Commons — Fangz — CC BY 4.0, via Wikimedia Commons
Vignette — The Underpowered "Negative" Trial

Vignette: A pilot trial (n = 38) of a new drug for diabetic gastroparesis shows symptom improvement, but the difference is not statistically significant (p = 0.09). The authors conclude "the drug is ineffective."

Best critique / next step: The study is likely underpowered — too small to detect a real effect — which raises the risk of a Type II (β) error (false negative). The correct move is to increase the sample size (or pool data) and repeat, not to declare the drug useless.

Key point: A non-significant p-value in a tiny trial reflects low power, not proven equivalence.

Type I vs Type II — Courtroom Classic

Set the null hypothesis = "the defendant is innocent."

  • Type I error (α) = convicting an innocent person → seeing a difference that isn't there (false positive)
  • Type II error (β) = letting the guilty go freemissing a difference that is real (false negative)
  • "Boy who cried wolf": Type I = shout "wolf!" with no wolf; Type II = stay silent while the wolf is real
  • Power = the ability to catch the real wolf — i.e., detect a true effect when it exists

Practice Biostatistics now

Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.