Confidence Intervals & Statistical Power
A Step 1-focused lesson on interpreting confidence intervals (crossing 0 for difference measures, 1 for ratios) and recognizing statistical power, Type I/II error, and the underpowered \"negative\" study.
Why This Matters on Step 1
Confidence intervals (CIs) and statistical power are the interpretation tools Step 1 loves, because they turn a bare p-value into a story. A confidence interval is the range of values likely to contain the true population parameter. A 95% CI means that if the study were repeated many times, 95% of the computed intervals would capture the truth — it is NOT "a 95% chance the true value lies in this one interval." Power is the mirror image: the probability that a study can actually detect an effect that is truly there. The boards almost never make you calculate. They make you interpret: does the interval cross the line of no effect, and was a "negative" trial simply too small to find a real difference?
- 95% CI = point estimate ± 1.96 × SEM, where SEM = SD/√n (99% CI uses 2.58; 90% uses 1.645)
- Higher confidence = WIDER interval (99% CI is wider than 95%); a wider CI is less precise
- A CI narrows with ↑ sample size, ↓ standard deviation, or ↓ confidence level
- Difference measures (mean difference, risk/rate difference): significant if the CI does NOT include 0
- Ratio measures (RR, OR, HR): significant if the CI does NOT include 1
- Two groups whose 95% CIs do not overlap differ significantly (p < 0.05); overlapping CIs are inconclusive
- A CI crossing the null carries the same verdict as p > 0.05 — but also shows effect size and precision
Vignette: An RCT of a new antihypertensive vs placebo reports a relative risk of stroke of 0.78 (95% CI 0.62–0.94).
Interpretation: This is a ratio measure, and the CI excludes 1, so the reduction is statistically significant (p < 0.05) — the drug lowers stroke risk.
Contrast: If instead RR were 0.85 (95% CI 0.70–1.05), the interval crosses 1 → not significant; you cannot conclude benefit despite the point estimate < 1.
Next step in reasoning: Confirm the significant effect is also clinically meaningful (magnitude, NNT), and check that the interval is narrow (precise) enough to trust rather than merely non-null.
Errors & Power at a Glance
| Concept | Meaning | Typical value |
|---|---|---|
| Type I error (α) | False positive — reject a TRUE null (see a difference that isn't there) | 0.05 |
| Type II error (β) | False negative — fail to reject a FALSE null (miss a real difference) | 0.20 |
| Power (1 − β) | Correctly detect a true effect that exists | ≥ 0.80 desired |
| p-value | Chance of data this extreme if the null were true | reject null if < α |
- Power = 1 − β = probability of correctly rejecting a false null (detecting a true effect)
- Increase power by: ↑ sample size (most controllable), ↑ effect size, ↑ α, ↓ variability (SD), one-tailed test
- Sample size is the lever investigators actually control — the usual board answer to "how do you boost power?"
- An underpowered study (small n) risks a Type II (β) error: a false "no difference"
- For a fixed n, α and β trade off: lowering α (fewer false positives) raises β (more false negatives)
- "Absence of evidence is not evidence of absence" — a non-significant result ≠ proof of no effect

Vignette: A pilot trial (n = 38) of a new drug for diabetic gastroparesis shows symptom improvement, but the difference is not statistically significant (p = 0.09). The authors conclude "the drug is ineffective."
Best critique / next step: The study is likely underpowered — too small to detect a real effect — which raises the risk of a Type II (β) error (false negative). The correct move is to increase the sample size (or pool data) and repeat, not to declare the drug useless.
Key point: A non-significant p-value in a tiny trial reflects low power, not proven equivalence.
Set the null hypothesis = "the defendant is innocent."
- Type I error (α) = convicting an innocent person → seeing a difference that isn't there (false positive)
- Type II error (β) = letting the guilty go free → missing a difference that is real (false negative)
- "Boy who cried wolf": Type I = shout "wolf!" with no wolf; Type II = stay silent while the wolf is real
- Power = the ability to catch the real wolf — i.e., detect a true effect when it exists
Practice Biostatistics now
Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.