Cohort & Case-Control Studies
A board-focused breakdown of the two core observational designs — cohort (group by exposure → follow forward → relative risk) versus case-control (group by disease → look backward → odds ratio) — with 2×2 setup, worked OR/RR computations, and the biases Step 1 loves to test.
The Core Distinction
Cohort and case-control studies are observational — the investigator watches and never assigns exposure (assigning exposure = a trial). They differ in the direction of reasoning.
A cohort study starts with exposure: it takes exposed and unexposed people who are disease-free and follows them forward to see who develops the outcome — measuring incidence and yielding a relative risk (RR).
A case-control study starts with the outcome: it takes people who already have the disease (cases) plus disease-free controls and looks backward to compare prior exposures — yielding an odds ratio (OR).
Boards test the pairing relentlessly: disease first → case-control → OR; exposure first → cohort → RR. Recognize the design from the first sentence of the vignette, then pick the correct measure of association.
- Cohort = group by exposure, follow forward → incidence + relative risk (RR). Can be prospective or retrospective.
- Case-control = group by disease, look backward → odds ratio (OR). Retrospective by design.
- Rare disease → use case-control (don't wait years for cases to accrue).
- Rare exposure → use cohort (start by assembling the exposed).
- Cohort assesses multiple outcomes of one exposure; case-control assesses multiple exposures for one disease.
- OR ≈ RR when the disease is rare (rare-disease assumption).
- Case-control's weak spot = recall bias; prospective cohort's = loss to follow-up + cost/time.
- Neither proves causation alone — both are observational and vulnerable to confounding.
Cohort vs Case-Control
| Feature | Cohort | Case-control |
|---|---|---|
| Groups defined by | Exposure | Disease |
| Direction | Forward (exposure→outcome) | Backward (outcome→exposure) |
| Timing | Prospective or retrospective | Retrospective |
| Measure of association | Relative risk (RR) | Odds ratio (OR) |
| Gives incidence? | Yes | No |
| Best when the... | exposure is rare | disease is rare |
| Can study multiple... | outcomes | exposures |
| Main bias | Loss to follow-up (attrition) | Recall / selection |
| Cost & time | High (if prospective) | Low, fast |
The 2×2 Table & Formulas
Always set up exposure as rows, disease as columns:
| | Disease + | Disease − | | --- | --- | --- | | Exposed | a | b | | Unexposed | c | d |
Relative risk (cohort) compares risk across rows: RR = [a/(a+b)] ÷ [c/(c+d)] — incidence in exposed over incidence in unexposed.
Odds ratio (case-control) is the cross-product: OR = ad/bc.
You cannot compute true incidence or RR from a case-control study, because the investigator fixed how many cases and controls to enroll — only the OR is valid there.
Interpretation is identical for both: >1 = associated with ↑ disease, =1 = no association, <1 = protective. Always check the 95% CI: if it includes 1, the result is not statistically significant.
Vignette: Investigators enroll 100 patients with pancreatic cancer and 100 cancer-free controls and ask about prior heavy coffee intake. Among cases, 90 were heavy drinkers; among controls, 30 were. What is the measure of association and its value?
Design: grouped by disease, looking back at exposure → case-control → use the odds ratio.
2×2: a = 90, b = 30, c = 10, d = 70.
OR = ad/bc = (90×70)/(30×10) = 6300/300 = 21.
Interpretation: heavy coffee intake is strongly associated with pancreatic cancer in this sample — but suspect confounding (smokers drink more coffee) and recall bias. Classic trap: a large OR from a case-control study shows association, not causation.
Vignette: Starting in 1948, researchers enrolled 5,209 disease-free adults in Framingham and followed them for decades, recording who developed coronary heart disease by baseline blood pressure. Incidence in hypertensives = 10%; in normotensives = 2%. Which design and measure?
Design: a disease-free group defined by exposure (baseline BP), followed forward → prospective cohort → relative risk.
RR = 10% / 2% = 5 — hypertensives had 5× the risk of CHD.
Next-step thinking: a prospective cohort minimizes recall bias (exposure is recorded before the outcome), but its weaknesses are cost, long duration, and loss to follow-up. Framingham is the board's prototype prospective cohort.
"Rare Disease → Case-Control; Rare Exposure → Cohort." Choose the design that starts with the rare thing, so cases (or exposed subjects) aren't impossibly slow to accumulate.
- Case-Control → begins with the outcome, looks back, computes the Odds Ratio (retrospective).
- Cohort → follows forward, computes the Relative Risk (measures true incidence).
- OR ≈ RR only when the disease is rare — the rare-disease assumption.
- Recall bias — differential memory of past exposure; the classic case-control flaw (cases over-report). Reduce with records/registries.
- Selection bias — cases and controls not comparable; Berkson bias (hospital-based controls) is the classic subtype.
- Attrition / loss to follow-up — the prospective cohort threat; differential dropout distorts the RR.
- Confounding — a third variable (e.g., smoking) linked to both exposure and outcome. Control in design (restriction, matching; randomization only in trials) or in analysis (stratification, multivariable regression).
- Latency — a long exposure-to-disease lag makes prospective cohorts impractical, favoring case-control / retrospective designs.
- Table orientation — swapping a single axis (exposed↔unexposed or diseased↔not) inverts RR/OR to its reciprocal, turning a risk factor into apparent protection; fix exposure = rows, disease = columns before computing.
Practice Research Methods now
Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.