Validity & Reliability
A Step 1 high-yield lesson distinguishing validity (accuracy; freedom from systematic error/bias) from reliability (precision; freedom from random error), anchored by the dartboard model and next-best-step vignettes on instrument calibration and inter-rater agreement.
Core framework: two independent properties
On the boards, validity and reliability describe two separate properties of any measurement or study.
Validity = accuracy — how close a measurement sits to the true value; it reflects freedom from systematic error (bias).
Reliability = precision — how reproducible a measurement is on repeat testing; it reflects freedom from random error.
The classic image is a dartboard: reliability is how tightly the darts cluster, validity is whether that cluster sits on the bullseye. The two are independent — a test can be highly reliable yet invalid (consistently wrong in the same direction), so precision never guarantees truth. The single highest-yield trap: increasing sample size shrinks random error (narrows the confidence interval) but does NOT fix bias (validity) — you cannot average your way out of a systematic error.
- Validity (accuracy): closeness to the true value; degraded by systematic error / bias; NOT fixed by a bigger sample — fix with better design, randomization, blinding, instrument calibration, and gold-standard comparison.
- Reliability (precision): reproducibility; degraded by random error; improved by standardizing technique and averaging repeated measures. ↑ sample size narrows the confidence interval (↑ precision of the estimate) but never removes bias.
- Random error → wider confidence interval + ↓ statistical power (unpredictable scatter around the truth).
- Systematic error (bias) → shifts the estimate away from truth in a predictable direction.
- Internal validity: the observed effect is the true effect within the study (free of bias/confounding) — maximized by randomization + blinding.
- External validity (generalizability): results apply to populations outside the study sample.
- Kappa (κ): inter-rater agreement beyond chance (κ = 1 perfect, κ = 0 chance-level, <0 worse than chance).
- Cronbach's α: internal consistency — do the items measure a single construct?
Validity vs. Reliability
| Feature | Validity | Reliability |
|---|---|---|
| Synonym | Accuracy / trueness | Precision / reproducibility |
| Error it removes | Systematic error (bias) | Random error |
| Dartboard | On the bullseye | Tightly clustered |
| Reduced by ↑ sample size? | No (bias persists) | Yes — narrows the CI of the estimate |
| Improved by | Randomization, blinding, calibration, gold-standard | Repeated/averaged measures, standardized protocol |
| Key metrics | Criterion, content, construct validity | Kappa (κ), Cronbach's α, test–retest |
Validity subtypes:
- Criterion: agreement with a gold standard — concurrent (measured now) vs predictive (predicts a future outcome).
- Content: the test samples the entire domain of the concept.
- Construct: the test measures the intended theoretical construct.
Reliability subtypes:
- Test–retest: same result on repeat administration over time.
- Inter-rater (inter-observer): different raters agree → quantified by kappa (κ).
- Internal consistency: items correlate with one another → Cronbach's α.
Picture darts on a target:
- Reliable + Valid = tight cluster on the bullseye → ideal.
- Reliable, NOT valid = tight cluster off the bullseye → systematic error / bias.
- Valid on average, NOT reliable = darts scattered around the bullseye → random error.
- Neither = scattered and off-center.
Honest cue: Reliable = Repeatable (same answer on every retest); Valid = hits the true target (accurate). So if repeat readings all agree but sit off-center, the test is precise but not accurate — a bias you cannot average away.
Vignette: A researcher weighs volunteers on a new digital scale. Repeat weighings of the same person give nearly identical readings, but every reading is exactly 2 kg higher than a calibrated reference scale.
Interpretation: The scale is highly reliable (precise) but not valid (inaccurate) — a consistent, directional offset is a systematic error.
Next best step: Recalibrate / re-zero the instrument against the reference standard. Averaging more measurements or enlarging the sample will not correct a systematic offset — that only tightens precision, which is already fine.
Vignette: Two radiologists independently read 100 chest CTs for pulmonary nodules; their reads agree only slightly more than expected by chance (κ = 0.15).
Interpretation: Poor inter-rater (inter-observer) reliability — near-random disagreement between observers — not a problem with CT accuracy itself.
Next best step: Improve reliability with a standardized reading protocol, explicit diagnostic criteria, and rater training, then re-measure κ.
Contrast: If a diet-drug trial enrolled only healthy 25-year-olds, applying its results to elderly patients is a failure of external validity (generalizability) — not reliability.
Practice Research Methods now
Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.