Skip to content
All lessons
Cross-cutting · Research Methods

Evidence-Based Medicine Principles

A boards-focused EBM lesson covering the exam's four core tasks — choosing study designs/statistical tests, interpreting sensitivity/specificity/PPV, computing RR/OR/ARR/NNT, and identifying bias — anchored in the 2×2 table with next-best-step reasoning.

13 min readHigh yield

What the boards actually test

Evidence-based medicine (EBM) integrates the best research evidence, clinical expertise, and patient values. Boards rarely ask for definitions — they test the mechanics: appraising study design, interpreting test performance, quantifying treatment effect, and spotting bias. Frame the clinical question with PICO (Patient, Intervention, Comparison, Outcome).

Know the hierarchy of evidence: systematic review / meta-analysis of RCTs > single RCT > cohort > case-control > cross-sectional > case series/report > expert opinion. Randomization is the single most powerful tool because it balances known and unknown confounders — nothing else does.

On exam day almost every EBM item is one of four flavors:

  1. Pick the study design or statistical test.
  2. Compute or interpret sensitivity / specificity / PPV / NPV.
  3. Compute RR / OR / ARR / NNT.
  4. Name the bias and how to fix it.

Master the 2×2 table and a handful of formulas and you own this topic.

Diagnostic test metrics

Build the 2×2 (disease across the top, test result down the side): TP, FP, FN, TN.

  • Sensitivity = TP/(TP+FN) — detects disease. High Sens, a Negative rules out (SnNout)
  • Specificity = TN/(TN+FP). High Spec, a Positive rules in (SpPin)
  • Sensitivity & specificity are intrinsic to the test — they do NOT change with prevalence
  • PPV = TP/(TP+FP) — rises with prevalence
  • NPV = TN/(TN+FN) — falls with prevalence
  • Low-prevalence screening → many false positives → low PPV (classic tested point)
  • LR+ = Sens/(1−Spec); LR− = (1−Sens)/Spec. LR+ >10 or LR− <0.1 = strong test
  • Lower the cutoff → ↑sensitivity (fewer FN, good screening test); raise the cutoff → ↑specificity (fewer FP, good confirmatory test)

Study designs at a glance

DesignDirectionMeasureBest for
RCTProspective, randomizedRR / HRCausation; gold standard for therapy
CohortExposure → outcomeRelative Risk (RR)Incidence; rare exposures; multiple outcomes
Case-controlOutcome → exposure (retrospective)Odds Ratio (OR)Rare diseases; cheap, fast
Cross-sectionalSnapshot in timePrevalence, OR"What's happening now"; prevalence
Meta-analysisPools prior studiesPooled effect↑ statistical power; top of hierarchy
Diagram showing true positives, false positives, false negatives, and true negatives with the formulas for sensitivity, specificity, PPV, and NPV
Every diagnostic-test metric falls out of the four cells: sensitivity and specificity are intrinsic to the test, while PPV and NPV shift with prevalence. · Wikimedia Commons — FeanDoe — CC BY-SA 4.0, via Wikimedia Commons
Vignette — prevalence drives PPV

Vignette: A company markets an HIV antibody screen with 99.9% sensitivity and 99.5% specificity. Deployed to low-risk blood donors (prevalence ≈ 1/10,000), the lab is flooded with positive results, yet nearly all flagged donors are uninfected on confirmatory testing.

Why? In a low-prevalence population, even a highly specific test yields a low PPV — false positives vastly outnumber true positives (here PPV ≈ 2%). Note the test's sensitivity and specificity are unchanged; only PPV shifted, driven by prevalence.

Next best step: Confirm every reactive screen with a confirmatory assay (HIV-1/2 antibody differentiation immunoassay) before disclosing a diagnosis. Never act on a single positive screening test in a low-prevalence setting.

Exam trap: If asked which parameter changed when the same test moves from a high-risk clinic to general screening, the answer is PPV (↓) and NPV (↑)not sensitivity or specificity.

Measures of effect & significance
  • Relative Risk (RR) = risk_exposed / risk_unexposed (cohort/RCT). RR =1 no effect, >1 harm, <1 protective
  • Odds Ratio (OR) = (a×d)/(b×c) (case-control); ≈ RR when disease is rare
  • Absolute Risk Reduction (ARR) = risk_control − risk_treated
  • Relative Risk Reduction (RRR) = 1 − RR = ARR / risk_control
  • Number Needed to Treat (NNT) = 1/ARR (round up); NNH = 1/absolute risk increase
  • Hazard Ratio (HR) = relative event rate over time (Cox/survival analysis)
  • Attributable risk % (in exposed) = (RR − 1)/RR

Confidence intervals: a ratio (RR/OR/HR) whose 95% CI crosses 1, or a difference of means whose CI crosses 0, is NOT statistically significant. Narrower CI = more precision (larger n).

High-yield biases and their fixes

BiasClassic scenarioFix
Selection (Berkson)Hospital-based controls not representativePopulation-based controls
RecallCases recall exposures better (case-control)Prospective design; use records
Lead-timeScreening detects disease earlier → apparent "longer survival," same death dateCompare mortality, not survival time
Length-timeScreening over-catches indolent/slow tumorsRandomized screening trial
ConfoundingCoffee–cancer link actually driven by smokingRandomize, restrict, match, stratify, regression
HawthorneSubjects change behavior when observedControl group
Observer/measurementAssessor knows the assignmentBlinding; objective criteria
Vignette — confounding vs effect modification

Vignette: A cohort finds coffee drinkers have the rate of lung cancer. After adjusting for cigarette smoking, the association vanishes (RR → 1.0). In a separate analysis, an occupational carcinogen's effect on lung cancer is much larger in smokers than non-smokers, and this stratum difference persists after adjustment.

First finding = confounding. Smoking is associated with both coffee and lung cancer and lies on no causal path from coffee → cancer; adjustment removes the spurious link.

Second finding = effect modification (interaction). The true effect genuinely differs by stratumreport stratum-specific estimates; do NOT "adjust it away" or report a single pooled number.

Next best step for a confounder: control it — randomization, restriction, matching (design), or stratification / multivariable regression (analysis). Randomization handles unknown confounders too.

Classics worth memorizing
  • SpPin — a highly Specific test, when Positive, rules in.
  • SnNout — a highly Sensitive test, when Negative, rules out.
  • Type I error (α) = false positive — you "cry wolf" / convict the innocent — reject a true null. Its rate is the significance level α (typically 0.05).
  • Type II error (β) = false negative — you "miss the wolf." Power = 1 − β, raised by larger sample size, larger effect size, and higher α.
  • Picking the test: compare means of 2 groups → t-test; 3+ groups → ANOVA; categorical/proportions → Chi-square ("Chi = chategorical").

Practice Research Methods now

Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.