Study Designs, Bias & Confounding
A board-focused walkthrough of the core epidemiology triad: matching each study design to its correct measure of association (OR, RR, prevalence), recognizing the classic biases, and distinguishing confounding from effect modification.
Framing the concept
Epidemiology questions on the boards almost always test three linked skills: matching a study design to the correct measure of association, recognizing the specific bias a scenario introduces, and distinguishing confounding from effect modification. Study designs fall into two families: observational (case-control, cohort, cross-sectional — the investigator only watches) and experimental (the randomized controlled trial — the investigator assigns the intervention). The design you are handed dictates the math you are allowed to do: a case-control study can only yield an odds ratio, a cohort study yields relative risk, and a cross-sectional study yields prevalence. Get the design right first, and the correct answer usually follows.
Study designs at a glance
| Study design | Type | What it does | Key measure |
|---|---|---|---|
| Case-control | Observational, retrospective | Starts with disease (cases vs controls), looks back for exposure | Odds ratio (OR) |
| Cohort | Observational, prospective or retrospective | Starts with exposure, follows forward for disease | Relative risk (RR); incidence |
| Cross-sectional | Observational, snapshot | Measures exposure and disease at one point in time | Prevalence; point-in-time association |
| RCT | Experimental | Randomly assigns intervention vs control | RR, ARR, NNT |
| Case report / series | Descriptive | Describes patients, no comparison group | None (hypothesis-generating) |
- Case-control -> OR; best for rare diseases and long latency; most prone to recall and selection bias.
- Cohort -> RR; best for rare exposures; the only observational design that measures incidence. Prospective cohorts are threatened by loss to follow-up.
- OR approximates RR only when the disease is rare.
- Cross-sectional measures prevalence and association but cannot establish temporality -> weak for causality.
- Randomization is the feature of the RCT that controls confounding — including unknown/unmeasured confounders. A crossover trial makes each patient their own control.
- Twin concordance (monozygotic vs dizygotic) and adoption studies both estimate heritability vs environmental contribution.
- Among the Bradford Hill criteria for causation, only temporality (cause precedes effect) is strictly required.
Recognize the design from the enrollment sentence, then pick the measure:
- *"Investigators identify patients with pancreatic cancer and a comparison group without, then ask about prior coffee intake"* -> case-control -> report an OR = ad/bc.
- *"Investigators enroll smokers and nonsmokers and follow them for 10 years for lung cancer"* -> cohort -> report RR = [a/(a+b)] / [c/(c+d)].
- *"A survey of a town on a single day measures both hypertension and BMI"* -> cross-sectional -> report prevalence.
Next step stems usually ask you to compute the right value from a 2x2 table, or to name why a design cannot yield RR (answer: sampling was done on the outcome — the investigator fixed the number of cases — so baseline disease risk/incidence is unknown).
Bias — systematic, not random
Bias is systematic error that pushes a study's result away from the truth and threatens its validity (accuracy). Crucially, unlike random error (which threatens precision and shrinks with a larger sample), bias is NOT reduced by increasing sample size — a big study can be confidently wrong. Bias is introduced during design, data collection, or analysis, and clusters into two big families tested on the boards: selection bias (who gets into or stays in the study) and information / measurement bias (how exposure and outcome are recorded).
High-yield biases
| Bias | Mechanism | Classic clue | How to reduce |
|---|---|---|---|
| Selection | Non-random entry/exit of subjects | Berkson (hospital-based controls), healthy-worker effect, loss to follow-up | Randomization; comparable groups |
| Recall | Cases remember past exposures differently | Retrospective case-control of birth defects | Use records; multiple data sources |
| Hawthorne effect | Subjects change behavior because they know they are observed | Behavior improves just from being studied | Objective/blinded measures; control group |
| Observer-expectancy / Pygmalion | Researcher's belief influences assessment | Unblinded outcome rating | Blinding; placebo |
| Lead-time | Screening detects disease earlier | "Survival from diagnosis" longer, mortality unchanged | Measure survival from disease onset; compare mortality |
| Length-time | Screening over-samples slow, indolent disease | Screen-detected cancers look less aggressive | RCT with a mortality endpoint |
- Bias = systematic error (validity); random error = imprecision (precision). Only random error shrinks with larger n.
- Selection-bias buzzwords: Berkson bias, healthy-worker effect, non-response bias, attrition/loss to follow-up.
- Recall bias is the signature weakness of retrospective / case-control studies.
- Hawthorne = subjects change because observed; Pygmalion / observer-expectancy = the researcher's expectation biases the result — fixed by blinding.
- Lead-time and length-time bias both make screening look falsely beneficial; the unbiased endpoint is disease-specific mortality, not survival time.
- Publication bias: positive/significant studies are over-published -> inflates meta-analysis effect sizes (assessed with a funnel plot).
Confounding vs effect modification
A confounder is a variable associated with both the exposure and the outcome but not on the causal pathway, creating a spurious association — the classic example is the apparent alcohol-lung cancer link that is actually driven by smoking. Confounding is controlled at the design stage (randomization, restriction, matching, crossover) and at the analysis stage (stratification, multivariable regression). Effect modification is fundamentally different: here the exposure-outcome effect genuinely differs across strata of a third variable — it is a real biologic finding, not an error, so you report stratum-specific estimates and do NOT adjust it away. The classic example is oral contraceptives + smoking on arterial thrombosis (MI, ischemic stroke) risk — the effect of OCPs is far larger in smokers, which is why combined OCPs are contraindicated in women over 35 who smoke. The tell: with a confounder the crude estimate differs from the adjusted one but the strata resemble each other; with effect modification the stratum-specific estimates differ from each other.
- *"Crude analysis shows coffee is associated with pancreatic cancer, but after adjusting for smoking the association disappears"* -> smoking is a confounder.
- *"RR of myocardial infarction with OCP use is 2 in nonsmokers but ~20 in smokers over 35"* -> smoking is an effect modifier (interaction); report the two strata separately rather than pooling.
- Only randomization controls for UNKNOWN/unmeasured confounders — a favorite "why is the RCT superior" answer.
- Next step to handle confounding in an observational study: stratified analysis or multivariable regression (or matching/restriction at the design stage).
Practice Research Methods now
Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.