Qualitative & Survey Methods
A high-yield boards lesson on qualitative and survey research methods, contrasting qualitative (hypothesis-generating, thematic, purposive sampling) with survey approaches and drilling the board-favorite concepts of reliability vs. validity (precision vs. accuracy; random vs. systematic error) and the common self-report biases with next-best-step fixes.
Clinical research isn't only trials and 2x2 tables. Qualitative research and survey/questionnaire methods answer questions numbers alone can't — the why and how behind behavior, and the measurement of attitudes at scale. Qualitative work (interviews, focus groups, observation) explores lived experience and generates hypotheses; it does not establish causation or prevalence. Surveys quantify beliefs across populations but live or die by two properties the boards love: reliability (consistency) and validity (accuracy), plus the biases that corrupt self-report. Expect vignettes that describe a study and ask you to name the method, identify the bias introduced, or choose the fix (next best step).
- Qualitative = non-numeric data (words, observations); explores why/how; generates hypotheses, does not test them or establish prevalence/causation
- Methods: in-depth interviews, focus groups, direct observation/ethnography; small purposive (non-random) samples
- Analysis = identify recurring themes; data saturation = point where no new themes emerge -> stop collecting data
- Reliability = reproducibility/precision -> same result on repeat (test-retest, inter-rater, internal consistency = Cronbach's alpha); degraded by random error
- Validity = accuracy -> the tool measures what it intends to (content, criterion, construct); degraded by systematic error (bias)
- A test can be reliable but NOT valid (consistently wrong); reliability is necessary but not sufficient for validity — a tool can't be valid without first being reliable
- Likert scale = ordinal agreement scale (strongly disagree -> strongly agree); closed-ended responses
Qualitative vs. Quantitative (survey) research
| Feature | Qualitative | Quantitative / Survey |
|---|---|---|
| Data | Words, observations, themes | Numbers, scores, counts |
| Question answered | Why / how | How many / how much |
| Role | Hypothesis-generating | Hypothesis-testing |
| Typical methods | Interviews, focus groups, ethnography | Structured questionnaires, cross-sectional surveys |
| Sampling | Small, purposive (non-random) | Larger, ideally random |
| Analysis | Thematic coding, saturation | Statistics, prevalence, correlation |
| Generalizable? | Limited (not statistical) | Yes, if sample representative |
A researcher gives a new 10-item depression questionnaire to the same clinically stable patients twice, one week apart, and gets nearly identical scores each time. However, the scores correlate poorly with a gold-standard structured psychiatric interview.
Interpretation: The instrument has high reliability (reproducible; strong test-retest) but poor validity — it is consistent but not measuring true depression (a systematic error is present).
Next best step: Revise the items to improve criterion validity (agreement with the gold standard). Remember: a tool can be reliable yet invalid, but it can never be valid without being reliable.
The classic target analogy keeps the two straight:
- Reliability = precision -> darts land tightly clustered together (reproducible), even if off-center
- Validity = accuracy -> darts land on the bullseye (the true value)
- Tight cluster, off-center = reliable but NOT valid (systematic error/bias)
- Scattered around the bullseye = valid on average but not reliable (random error)
- Tight cluster on the bullseye = the goal: reliable AND valid
Subtypes of reliability and validity
| Property | Subtype | Meaning |
|---|---|---|
| Reliability | Test-retest | Same score on repeat over time |
| Inter-rater | Different observers agree | |
| Internal consistency | Items measure same construct (Cronbach's alpha) | |
| Validity | Content | Items cover the full concept |
| Criterion | Correlates with gold standard (concurrent = now; predictive = future) | |
| Construct | Measures the intended theoretical trait |
- Nonresponse bias — non-responders differ systematically from responders; a low response rate threatens generalizability (external validity) -> boost with reminders/incentives, shorter survey
- Volunteer / self-selection bias — people who opt in differ from the target population
- Recall bias — inaccurate memory of past exposures (differential; often worse in cases than controls)
- Social desirability bias — respondents give socially acceptable answers (underreport alcohol/drug use, overreport exercise) -> use anonymous, self-administered surveys
- Acquiescence (response) bias — tendency to agree regardless of item content
- Leading-question bias — wording steers the answer -> use neutral wording
- Hawthorne effect — subjects change behavior because they know they are being observed
A mailed survey on illicit drug use achieves only a 25% response rate; drug users are less likely to reply, and those who do respond underreport their use.
Biases present: the low response rate with systematically different non-responders -> nonresponse bias; underreporting a stigmatized behavior -> social desirability bias.
Next best step: Increase the response rate (reminders, incentives, shorter survey) and minimize social-desirability effects by making the questionnaire anonymous and self-administered rather than interviewer-conducted.
Practice Research Methods now
Board-style questions, spaced-repetition flashcards, and a Socratic AI tutor — free to start.