Epidemiology and Study Design
Contents (10)
What the topic is
- Epidemiology: the study of the distribution and determinants of disease in defined populations, and the application of that knowledge to disease control. The unit of analysis is the population, not the patient.
- Study design: the architecture that determines what a dataset can and cannot prove. Design dictates which effect measure is calculable (risk vs. odds), whether temporality can be established, and which biases are structurally possible.
Why it matters clinically
- Every guideline recommendation a resident follows is a design judgment. The USPSTF grades preventive services A through D and I based largely on the strength and directness of the underlying evidence, and under the Affordable Care Act, A and B recommendations must be covered by most insurers without cost sharing โ so design quality has direct financial and access consequences.
- The ACC/AHA framework pairs a Class of Recommendation (magnitude and certainty of benefit vs. risk) with a Level of Evidence, and since the 2015 methodology update the LOE tiers are subdivided: A (high-quality evidence from more than one RCT, or meta-analyses of high-quality RCTs), B-R (moderate-quality evidence from one or more RCTs), B-NR (moderate-quality nonrandomized/observational/registry data), C-LD (limited data, including observational studies with design limitations or physiologic studies), and C-EO (consensus expert opinion). The two axes are independent: a Class I/LOE C-EO recommendation and a Class I/LOE A recommendation carry the same directive force but very different empirical backing.
- The GRADE framework, used widely by IDSA, ATS, and KDIGO, starts RCTs at high certainty and observational studies at low certainty, then adjusts for risk of bias, imprecision, inconsistency, indirectness, and publication bias.
Distribution worth recalling
- Observational designs (cohort, case-control, cross-sectional) vastly outnumber randomized trials in the published literature, because randomization is often unethical (tobacco, teratogens), impractical for rare outcomes, or too slow for long-latency disease.
- Cross-sectional (prevalence) surveys such as NHANES and BRFSS are the backbone of U.S. chronic disease surveillance; registry-based cohorts dominate cancer and cardiovascular outcome reporting.
Frequency measures
- Cumulative incidence (risk): new cases รท population at risk at baseline, over a defined interval. It is a proportion bounded 0โ1 and is meaningless without stating the time period.
- Incidence rate (incidence density): new cases รท person-time at risk. It is a rate with units of 1/time and no upper bound; it handles variable follow-up and loss to follow-up, which cumulative incidence cannot. Only a longitudinal design produces either measure.
- Prevalence: existing cases รท total population at one point. Prevalence โ incidence ร average disease duration. A therapy that prolongs life without curing (antiretrovirals in HIV) raises prevalence while incidence falls.
Effect measures
- Relative risk (RR) = risk in exposed รท risk in unexposed. Interpretable only when denominators of at-risk people are known โ cohort studies and RCTs.
- Odds ratio (OR) = ad/bc from the 2ร2 table. In case-control studies the investigator fixes the number of cases, so absolute risk is unknowable and only odds can be computed.
- Attributable risk (risk difference) = risk exposed โ risk unexposed. Answers "how much disease would disappear if exposure vanished" โ the public health number.
- Absolute risk reduction (ARR) and relative risk reduction (RRR = ARR รท control risk). RRR is the number drug advertisements quote because it looks larger; ARR determines whether the effect is clinically worth the harm.
- Hazard ratio: an instantaneous relative rate over follow-up time, the output of Cox regression and KaplanโMeierโbased analyses.
Test performance beyond sensitivity/specificity
- Likelihood ratio positive = sensitivity รท (1 โ specificity); LR negative = (1 โ sensitivity) รท specificity. LRs are prevalence-independent and convert pretest to posttest odds. LR+ above roughly 10 or LRโ below roughly 0.1 shifts probability substantially.
- ROC curve and AUC: plots sensitivity against 1 โ specificity across cut-points; AUC 0.5 is a coin flip, 1.0 is perfect discrimination.
Inference
- Type I error (ฮฑ): rejecting a true null โ a false positive finding. Type II error (ฮฒ): missing a real effect. Power = 1 โ ฮฒ, increased chiefly by larger sample size and larger effect size.
- Confidence interval: crossing 1 for a ratio measure (RR, OR, HR) or 0 for a difference measure means non-significance.
- Confounder vs. effect modifier: a confounder distorts the association and should be adjusted away; an effect modifier means the true effect genuinely differs by stratum and must be reported separately, never pooled.
Worked stem โ occupational cohort: Investigators follow 200 solvent-exposed workers and 400 unexposed workers for 10 years. Among the exposed, 40 develop peripheral neuropathy; among the unexposed, 20 do.
- Risk in exposed = 40/200 = 20%. Risk in unexposed = 20/400 = 5%.
- RR = 0.20 รท 0.05 = 4.0 โ exposed workers have four times the risk.
- Attributable risk = 20% โ 5% = 15%; the attributable risk percent = (RR โ 1)/RR = 75%, meaning three-quarters of disease in exposed workers is ascribable to the solvent.
- NNH = 1 รท 0.15 โ 7 workers exposed for one extra case of neuropathy.
- If a student mistakenly computes the OR = (40/160) รท (20/380) = 4.75, note that it overshoots the RR. This is the rare disease assumption failing: with 20% risk in the exposed, OR no longer approximates RR and always exaggerates away from 1.
Worked stem โ screening and lead-time: A new blood test detects pancreatic cancer earlier. Five-year survival rises from 6% to 15%, but median age at death is unchanged.
- The single best interpretation is lead-time bias: earlier diagnosis lengthens the interval between diagnosis and death without postponing death. The correct endpoint is disease-specific mortality, not survival from diagnosis โ the reason USPSTF screening recommendations are anchored to mortality outcomes.
- Its partner, length-time bias, is the tendency of any screening program to preferentially capture indolent, slow-growing tumors, inflating apparent screening benefit; its extreme form is overdiagnosis.
Worked stem โ trial analysis: In an RCT of a surgical versus medical strategy, 15% of the surgical arm never undergoes surgery. Analyzing only those who actually had surgery breaks randomization and reintroduces confounding by indication. The intention-to-treat analysis, required by CONSORT reporting standards, preserves baseline comparability and generally biases superiority trials toward the null. In non-inferiority trials the ITT analysis is not automatically conservative โ dropout and crossover blur the arms together and can falsely support non-inferiority โ so FDA guidance advises reporting both ITT and per-protocol analyses, with non-inferiority ideally supported by both.
- Recall bias is the case-control signature: mothers of infants with birth defects scrutinize their pregnancy exposures far more than mothers of healthy infants. Any retrospective self-reported exposure stem should trigger it. Minimize with medical records or biomarkers rather than interviews.
- Lead-time bias improves survival without improving mortality โ the single most tested screening trap. If the stem gives you 5-year survival after a new early-detection test and unchanged age at death, the answer is lead-time bias, not a real benefit.
- The rare disease assumption is the whole reason OR is taught: OR approximates RR only when the outcome is uncommon. When outcome risk is high, OR is always farther from 1 than RR โ a classic distractor in stems that report both.
- Randomization controls confounding, blinding controls information bias. These are not interchangeable. Randomization addresses unmeasured confounders, which is the one thing no statistical adjustment can do.
- Hawthorne effect: subjects change behavior because they know they are observed. Contrast with the placebo effect (response to inert intervention) and the observer-expectancy effect (the researcher influences the result) โ examiners love these three side by side.
- Healthy worker effect makes employed cohorts look healthier than the general population; the correct comparison group is other workers, not the general public. Its cousin, Neyman (prevalenceโincidence) bias, occurs when rapidly fatal cases die before they can be enrolled, so survivors misrepresent the disease.
- Ecological fallacy: inferring an individual-level association from group-level data. A stem comparing per-country fat consumption to per-country cancer rates cannot say anything about any individual patient.
- The single best next step when an association appears: check temporality and confounding before invoking causation. Apply the Bradford Hill considerations โ strength, consistency, temporality, biologic gradient, plausibility โ remembering that only temporality is strictly necessary.
- Study designs ranked by evidence quality: RCT > Cohort > Case-Control > Cross-sectional > Case reports
- Sensitivity = TP/(TP+FN) โ ability to rule OUT disease; Specificity = TN/(TN+FP) โ ability to rule IN disease
- Relative Risk (RR) used in prospective studies; Odds Ratio (OR) used in case-control studies
- Number Needed to Treat (NNT) = 1/ARR; Number Needed to Harm (NNH) = 1/ARR for adverse effect
- Positive Predictive Value (PPV) and Negative Predictive Value (NPV) depend on disease prevalence
Epidemiologic study design selection depends on research question, timeframe, and feasibility. Prospective studies (RCT, cohort) measure outcomes forward in time and calculate RR; retrospective studies (case-control) look backward and use OR. Validity (internal/external) and reliability (reproducibility) determine study quality. Confounding, selection bias, and information bias distort results. Incidence measures new cases; prevalence measures existing cases at a point in time.
"A researcher wants to know if a new vaccine prevents disease X. She randomly assigns 1,000 people to vaccine or placebo and follows them for 2 years."
โ RCT (gold standard, highest evidence)
"A case-control study found OR = 3.2 for smoking and lung cancer."
โ Compare exposed cases vs. unexposed controls retrospectively
"A screening test has sensitivity 95% and specificity 85% in a population with 1% prevalence."
โ PPV will be LOW (~6%) despite high sensitivity (low prevalence = false positives dominate)
| Bias Type | Definition | How to Minimize |
|---|---|---|
| Selection Bias | Systematic difference in who enrolls | Randomization, define inclusion criteria clearly |
| Information Bias | Systematic error in exposure/outcome measurement | Standardized data collection, blinding |
| Confounding | 3rd variable causes apparent association | Matching, stratification, multivariate analysis |
| Berkson's Bias | Hospital-based samples (non-representative) | Population-based studies |
Memory Aid โ SpPin/SnNout
- High Specificity โ Positive test rules IN disease
- High Sensitivity โ Negative test rules OUT disease
- Confusing RR vs. OR: RR = prospective studies only; OR approximates RR when disease is rare (<10%). High-prevalence diseases: RR โ OR
- PPV/NPV trap: A highly sensitive test with low prevalence = many false positives (low PPV). Must account for pretest probability; don't rely on sensitivity/specificity alone
- Study design mismatch: Using case-control for common diseases (inefficient) or RCT for rare outcomes (impractical). Match design to research question and feasibility
N/A โ Epidemiology is methodologic, not a disease state. Focus is on study design selection, bias reduction, and accurate interpretation of sensitivity/specificity, PPV/NPV, and effect measures (RR/OR/NNT) based on clinical context.
Exam Tip: USMLE heavily tests sensitivity/specificity/PPV/NPV relationships with prevalence, study design hierarchy, and identification of bias types. Know SnNout/SpPin cold.
Related topics
- Study Designs in Clinical ResearchPublic Health Sciences
- Study Design and Evidence LevelsPublic Health Sciences
- Bias and Confounding in ResearchPublic Health Sciences
- Epidemiology โ Study DesignsPublic Health Sciences
- Statistical Measures and BiasPublic Health Sciences
- Advance Directives and Surrogate Decision-MakingPublic Health Sciences