Study Designs in Clinical Research
Contents (4)
Design determines which measure of association can be calculated and which biases threaten the result — the two things questions test.
- Cross-sectional study: exposure and outcome measured at one point in time. Yields prevalence; cannot establish temporality, so it generates hypotheses rather than causal claims.
- Case–control study: starts with the outcome and looks backwards for exposure. Efficient for rare diseases and long latency; yields an odds ratio, and cannot yield incidence or relative risk. Vulnerable to recall and selection bias.
- Cohort study: starts with exposure and follows forward for outcome. Yields incidence and relative risk; good for rare exposures; prospective cohorts are costly and vulnerable to loss to follow-up.
- Randomised controlled trial: randomisation balances known and unknown confounders, which no observational design can do. Blinding limits observer and reporting bias. Limited by cost, generalisability and ethics.
- Crossover trial: each participant serves as their own control; watch for carryover effects.
- Meta-analysis pools studies for precision but inherits their flaws and is subject to publication bias.
- Case series and case reports describe without a comparison group and cannot measure association.
(Seed article — remaining sections to be written and reviewed.)
Frequency measures
- Prevalence: existing cases ÷ total population at one moment. The natural output of a cross-sectional study; inflated by anything that prolongs disease duration (better treatment that averts death but not disease).
- Incidence: new cases ÷ population at risk per unit time. Requires follow-up, so only cohorts and trials produce it. Conceptually, prevalence ≈ incidence × average disease duration.
Measures of association (2×2 table with a = exposed cases, b = exposed non-cases, c = unexposed cases, d = unexposed non-cases)
- Relative risk (RR) = [a/(a+b)] ÷ [c/(c+d)]. Interpretable only when the denominators are real populations — i.e., cohort or RCT. RR = 1 means no association.
- Odds ratio (OR) = ad/bc. In a case–control study the investigator fixes how many cases and controls to enrol, so risk denominators are artefacts of sampling; the exposure odds ratio survives that sampling and is the only valid effect measure. OR approximates RR when the outcome is uncommon (rule of thumb, roughly under 10%); with a common outcome the OR exaggerates the RR away from 1.
- Attributable risk / risk difference = risk in exposed − risk in unexposed. Absolute, not relative — the basis of ARR, NNT = 1/ARR, and NNH = 1/attributable risk increase.
- Attributable risk percent = (RR − 1)/RR: the fraction of disease among the exposed attributable to the exposure.
- Hazard ratio: ratio of instantaneous event rates from time-to-event (Cox) analysis; handles censoring and unequal follow-up, so it is the standard trial and cohort output.
Analysis and reporting conventions
- Intention-to-treat: analyse participants in the arm to which they were randomised regardless of adherence — preserves randomisation and gives the less biased, usually more conservative estimate; per-protocol analysis reintroduces confounding. Codified in the CONSORT statement for trials.
- Reporting frameworks: STROBE for observational studies, CONSORT for randomised trials, PRISMA 2020 for systematic reviews and meta-analyses, GRADE for rating certainty of evidence; ICMJE requires prospective trial registration (e.g., ClinicalTrials.gov) before enrolment.
Worked stem 1 — recognising the design from the sampling frame. "Investigators enrol 200 patients with newly diagnosed pancreatic cancer and 200 hospitalised patients without cancer, then ask both groups about prior smoking. Smoking is reported by 120 cases and 60 controls." Enrolment was by outcome, so this is a case–control study and the answer must be an odds ratio: OR = ad/bc = (120 × 140)/(80 × 60) = 3.5. A choice offering "relative risk 2.0" is the trap — no incidence exists here because the case:control ratio was chosen by the investigator.
Worked stem 2 — cohort arithmetic. 1,000 exposed workers yield 40 cases and 1,000 unexposed yield 10 cases over 10 years. Risks are 4% and 1%: RR = 4, attributable risk = 3%, attributable risk percent = (4 − 1)/4 = 75%. Relative measures argue for causation; the absolute difference tells you the public-health yield of removing the exposure.
Worked stem 3 — trial arithmetic. Event rate 20% on placebo versus 12% on drug: ARR = 8%, NNT = 1/0.08 = 12.5, reported as 13 (always round NNT up). Relative risk reduction = 40% — the same data made to sound larger, a favourite distractor in stems about pharmaceutical advertising.
Choosing the design
- Rare disease or long latency (mesothelioma, Creutzfeldt–Jakob): case–control.
- Rare or unusual exposure (an occupational solvent, a single drug lot): cohort.
- Burden of disease at a moment, or screening-program planning: cross-sectional, and note that USPSTF recommendations rest on the trial and cohort evidence such surveys cannot supply.
- Definitive efficacy claim: randomised trial, analysed by intention-to-treat per CONSORT, with prospective registration per ICMJE.
- Existing cohort with banked specimens: nested case–control — assays are done only on sampled controls, but exposure still precedes outcome.
- Sampling frame names the design: enrolled by outcome = case–control (odds ratio only); enrolled by exposure = cohort (incidence, RR); everything measured at once = cross-sectional (prevalence). This single question resolves most stems.
- Case–control studies cannot yield incidence, relative risk, or attributable risk — the most frequently tested distractor on the exam.
- OR ≈ RR only when the outcome is rare. With a common outcome the odds ratio overstates the effect; a stem describing a 40% event rate and reporting an OR is signalling exactly this.
- A better treatment that prolongs survival raises prevalence while incidence is unchanged (prevalence ≈ incidence × duration). Do not read rising prevalence as rising risk.
- NNT = 1/ARR, always rounded up, and relative risk reduction can look impressive while the ARR — and thus clinical value — is trivial.
- Intention-to-treat preserves randomisation and is the CONSORT-endorsed primary analysis; per-protocol analysis reintroduces the confounding randomisation was meant to erase. "Analyse only adherent patients" is never the best answer for a primary efficacy endpoint.
- Bias by design: recall bias and Berkson bias (hospital controls) belong to case–control; loss to follow-up and healthy worker effect to cohorts; Neyman (prevalence–incidence) bias to cross-sectional studies, which miss rapidly fatal or rapidly resolving disease; publication bias, detected as funnel-plot asymmetry, to meta-analysis.
- Use medical records rather than interviews, and blind the interviewer, to blunt recall bias.
- Matching cases and controls on a variable eliminates it as a confounder but also makes it unstudiable, and requires matched analysis — a classic single-best-answer trap.
- Only randomisation balances unknown confounders; stratification, matching, restriction, and multivariable regression handle measured confounders only.
Related topics
- Epidemiology and Study DesignPublic Health Sciences
- Study Design and Evidence LevelsPublic Health Sciences
- Statistical Measures and BiasPublic Health Sciences
- Advance Directives and Surrogate Decision-MakingPublic Health Sciences
- Bias and Confounding in ResearchPublic Health Sciences
- BiostatisticsPublic Health Sciences