Skip to content
All lessons
Cross-cutting · Research Methods

Validity & Reliability

A Step 1 high-yield lesson distinguishing validity (accuracy; freedom from systematic error/bias) from reliability (precision; freedom from random error), anchored by the dartboard model and next-best-step vignettes on instrument calibration and inter-rater agreement.

8 min readHigh yield

Core framework: two independent properties

On the boards, validity and reliability describe two separate properties of any measurement or study.

Validity = accuracy — how close a measurement sits to the true value; it reflects freedom from systematic error (bias).

Reliability = precision — how reproducible a measurement is on repeat testing; it reflects freedom from random error.

The classic image is a dartboard: reliability is how tightly the darts cluster, validity is whether that cluster sits on the bullseye. The two are independent — a test can be highly reliable yet invalid (consistently wrong in the same direction), so precision never guarantees truth. The single highest-yield trap: increasing sample size shrinks random error (narrows the confidence interval) but does NOT fix bias (validity) — you cannot average your way out of a systematic error.

Must-know facts
  • Validity (accuracy): closeness to the true value; degraded by systematic error / bias; NOT fixed by a bigger sample — fix with better design, randomization, blinding, instrument calibration, and gold-standard comparison.
  • Reliability (precision): reproducibility; degraded by random error; improved by standardizing technique and averaging repeated measures. ↑ sample size narrows the confidence interval (↑ precision of the estimate) but never removes bias.
  • Random error → wider confidence interval + ↓ statistical power (unpredictable scatter around the truth).
  • Systematic error (bias) → shifts the estimate away from truth in a predictable direction.
  • Internal validity: the observed effect is the true effect within the study (free of bias/confounding) — maximized by randomization + blinding.
  • External validity (generalizability): results apply to populations outside the study sample.
  • Kappa (κ): inter-rater agreement beyond chance (κ = 1 perfect, κ = 0 chance-level, <0 worse than chance).
  • Cronbach's α: internal consistency — do the items measure a single construct?

Validity vs. Reliability

FeatureValidityReliability
SynonymAccuracy / truenessPrecision / reproducibility
Error it removesSystematic error (bias)Random error
DartboardOn the bullseyeTightly clustered
Reduced by ↑ sample size?No (bias persists)Yes — narrows the CI of the estimate
Improved byRandomization, blinding, calibration, gold-standardRepeated/averaged measures, standardized protocol
Key metricsCriterion, content, construct validityKappa (κ), Cronbach's α, test–retest
Subtypes to recognize

Validity subtypes:

  • Criterion: agreement with a gold standard — concurrent (measured now) vs predictive (predicts a future outcome).
  • Content: the test samples the entire domain of the concept.
  • Construct: the test measures the intended theoretical construct.

Reliability subtypes:

  • Test–retest: same result on repeat administration over time.
  • Inter-rater (inter-observer): different raters agree → quantified by kappa (κ).
  • Internal consistency: items correlate with one another → Cronbach's α.
The dartboard (classic)

Picture darts on a target:

  • Reliable + Valid = tight cluster on the bullseye → ideal.
  • Reliable, NOT valid = tight cluster off the bullseye → systematic error / bias.
  • Valid on average, NOT reliable = darts scattered around the bullseye → random error.
  • Neither = scattered and off-center.

Honest cue: Reliable = Repeatable (same answer on every retest); Valid = hits the true target (accurate). So if repeat readings all agree but sit off-center, the test is precise but not accurate — a bias you cannot average away.

Vignette 1 — the miscalibrated scale

Vignette: A researcher weighs volunteers on a new digital scale. Repeat weighings of the same person give nearly identical readings, but every reading is exactly 2 kg higher than a calibrated reference scale.

Interpretation: The scale is highly reliable (precise) but not valid (inaccurate) — a consistent, directional offset is a systematic error.

Next best step: Recalibrate / re-zero the instrument against the reference standard. Averaging more measurements or enlarging the sample will not correct a systematic offset — that only tightens precision, which is already fine.

Vignette 2 — disagreeing radiologists

Vignette: Two radiologists independently read 100 chest CTs for pulmonary nodules; their reads agree only slightly more than expected by chance (κ = 0.15).

Interpretation: Poor inter-rater (inter-observer) reliability — near-random disagreement between observers — not a problem with CT accuracy itself.

Next best step: Improve reliability with a standardized reading protocol, explicit diagnostic criteria, and rater training, then re-measure κ.

Contrast: If a diet-drug trial enrolled only healthy 25-year-olds, applying its results to elderly patients is a failure of external validity (generalizability) — not reliability.

Practice Research Methods now

Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.