Skip to content
All lessons
Cross-cutting · Research Methods

Cohort & Case-Control Studies

A board-focused breakdown of the two core observational designs — cohort (group by exposure → follow forward → relative risk) versus case-control (group by disease → look backward → odds ratio) — with 2×2 setup, worked OR/RR computations, and the biases Step 1 loves to test.

11 min readHigh yield

The Core Distinction

Cohort and case-control studies are observational — the investigator watches and never assigns exposure (assigning exposure = a trial). They differ in the direction of reasoning.

A cohort study starts with exposure: it takes exposed and unexposed people who are disease-free and follows them forward to see who develops the outcome — measuring incidence and yielding a relative risk (RR).

A case-control study starts with the outcome: it takes people who already have the disease (cases) plus disease-free controls and looks backward to compare prior exposures — yielding an odds ratio (OR).

Boards test the pairing relentlessly: disease first → case-control → OR; exposure first → cohort → RR. Recognize the design from the first sentence of the vignette, then pick the correct measure of association.

Must-Know Facts
  • Cohort = group by exposure, follow forward → incidence + relative risk (RR). Can be prospective or retrospective.
  • Case-control = group by disease, look backward → odds ratio (OR). Retrospective by design.
  • Rare disease → use case-control (don't wait years for cases to accrue).
  • Rare exposure → use cohort (start by assembling the exposed).
  • Cohort assesses multiple outcomes of one exposure; case-control assesses multiple exposures for one disease.
  • OR ≈ RR when the disease is rare (rare-disease assumption).
  • Case-control's weak spot = recall bias; prospective cohort's = loss to follow-up + cost/time.
  • Neither proves causation alone — both are observational and vulnerable to confounding.

Cohort vs Case-Control

FeatureCohortCase-control
Groups defined byExposureDisease
DirectionForward (exposure→outcome)Backward (outcome→exposure)
TimingProspective or retrospectiveRetrospective
Measure of associationRelative risk (RR)Odds ratio (OR)
Gives incidence?YesNo
Best when the...exposure is raredisease is rare
Can study multiple...outcomesexposures
Main biasLoss to follow-up (attrition)Recall / selection
Cost & timeHigh (if prospective)Low, fast

The 2×2 Table & Formulas

Always set up exposure as rows, disease as columns:

| | Disease + | Disease − | | --- | --- | --- | | Exposed | a | b | | Unexposed | c | d |

Relative risk (cohort) compares risk across rows: RR = [a/(a+b)] ÷ [c/(c+d)] — incidence in exposed over incidence in unexposed.

Odds ratio (case-control) is the cross-product: OR = ad/bc.

You cannot compute true incidence or RR from a case-control study, because the investigator fixed how many cases and controls to enroll — only the OR is valid there.

Interpretation is identical for both: >1 = associated with ↑ disease, =1 = no association, <1 = protective. Always check the 95% CI: if it includes 1, the result is not statistically significant.

Vignette — Compute the OR

Vignette: Investigators enroll 100 patients with pancreatic cancer and 100 cancer-free controls and ask about prior heavy coffee intake. Among cases, 90 were heavy drinkers; among controls, 30 were. What is the measure of association and its value?

Design: grouped by disease, looking back at exposure → case-control → use the odds ratio.

2×2: a = 90, b = 30, c = 10, d = 70.

OR = ad/bc = (90×70)/(30×10) = 6300/300 = 21.

Interpretation: heavy coffee intake is strongly associated with pancreatic cancer in this sample — but suspect confounding (smokers drink more coffee) and recall bias. Classic trap: a large OR from a case-control study shows association, not causation.

Vignette — Identify the Design

Vignette: Starting in 1948, researchers enrolled 5,209 disease-free adults in Framingham and followed them for decades, recording who developed coronary heart disease by baseline blood pressure. Incidence in hypertensives = 10%; in normotensives = 2%. Which design and measure?

Design: a disease-free group defined by exposure (baseline BP), followed forwardprospective cohortrelative risk.

RR = 10% / 2% = 5 — hypertensives had the risk of CHD.

Next-step thinking: a prospective cohort minimizes recall bias (exposure is recorded before the outcome), but its weaknesses are cost, long duration, and loss to follow-up. Framingham is the board's prototype prospective cohort.

Classic Memory Aids

"Rare Disease → Case-Control; Rare Exposure → Cohort." Choose the design that starts with the rare thing, so cases (or exposed subjects) aren't impossibly slow to accumulate.

  • Case-Control → begins with the outcome, looks back, computes the Odds Ratio (retrospective).
  • Cohort → follows forward, computes the Relative Risk (measures true incidence).
  • OR ≈ RR only when the disease is rare — the rare-disease assumption.
Biases & Pitfalls the USMLE Loves
  • Recall bias — differential memory of past exposure; the classic case-control flaw (cases over-report). Reduce with records/registries.
  • Selection bias — cases and controls not comparable; Berkson bias (hospital-based controls) is the classic subtype.
  • Attrition / loss to follow-up — the prospective cohort threat; differential dropout distorts the RR.
  • Confounding — a third variable (e.g., smoking) linked to both exposure and outcome. Control in design (restriction, matching; randomization only in trials) or in analysis (stratification, multivariable regression).
  • Latency — a long exposure-to-disease lag makes prospective cohorts impractical, favoring case-control / retrospective designs.
  • Table orientation — swapping a single axis (exposed↔unexposed or diseased↔not) inverts RR/OR to its reciprocal, turning a risk factor into apparent protection; fix exposure = rows, disease = columns before computing.

Practice Research Methods now

Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.