Study design in pharmacoepidemiology

Reading File
Finding Sources
Searching PubMed

pharmacoepidemiology study design observational

Finding Sources
Finding Sources
Finding Sources
Finding Sources
Finding Sources
Reading File
Finding Sources
Reading File
Reading File
Reading File
Reading File
Reading File
Reading File
Reading File
Reading File
Reading File
Reading File
I now have comprehensive material from the textbooks. Let me compile the full, well-organized answer.

Study Design in Pharmacoepidemiology

Pharmacoepidemiology is the study of drug use and effects - both beneficial and harmful - in large populations under real-world conditions. It bridges clinical pharmacology and epidemiology. Because randomized controlled trials (RCTs) cannot answer every question about drug safety and effectiveness (especially for rare outcomes, long-term harms, or special populations), pharmacoepidemiology relies heavily on observational designs applied to real-world data.

1. Classification of Study Designs

Study designs are broadly classified as:
CategoryExamples
ExperimentalRCT, pragmatic trial, cluster RCT
Observational - longitudinalCohort (prospective), case-control (retrospective)
Observational - cross-sectionalPrevalence surveys, drug utilization surveys
Special/hybridNested case-control, case-cohort, case-crossover, ecological
The first distinction is between longitudinal and cross-sectional studies. Longitudinal studies observe change over time; cross-sectional studies describe phenomena at a single point in time (e.g., reporting drug utilization frequencies).
  • Barash's Clinical Anesthesia, p. 501

2. Cohort Studies

Definition

A cohort study assembles a group of individuals and classifies them by exposure status (drug user vs. non-user), then follows them forward to observe outcomes.

Types

  • Prospective cohort: Exposure and follow-up occur in real time. Expensive but high quality.
  • Retrospective cohort: Uses historical records. Faster and cheaper.
  • Closed cohort: Fixed membership, all enrolled at once (like a graduating class).
  • Open (dynamic) cohort: Participants may enter and exit over time (more common with administrative databases).

Key features

  • Classifies by exposure first, then observes outcomes - allows calculation of incidence rates and risk ratios.
  • Can study multiple outcomes simultaneously from a single exposed group.
  • Best for common outcomes and when long follow-up is affordable.
  • Subject to loss to follow-up (censoring) - person-time analysis is used to account for this.

Measure of effect

Risk ratio (RR), incidence rate ratio (IRR), hazard ratio (HR)
  • Firestein & Kelley's Rheumatology, p. 611

3. Case-Control Studies

Definition

Individuals are sampled based on outcome status: those with the disease (cases) and those without (controls). Exposure history is then assessed and compared between the two groups.

Types

SubtypeDescription
Population-basedControls drawn from the general population
Hospital-basedControls drawn from patients with other conditions
Nested case-controlCases and controls are drawn from within a defined cohort - the most rigorous form
Case-cohortControls (the subcohort) are randomly sampled at baseline from the full cohort before outcomes occur

Key features

  • Efficient for rare outcomes - you only need to enroll people who have had the outcome, plus a sample of controls.
  • Can study multiple exposures for a single outcome.
  • Retrospective nature increases risk of recall bias (differential misclassification of exposure).
  • Controls must be selected such that: (1) they would have been captured as a case if they had developed the outcome, and (2) selection is independent of exposure.
  • Risk set sampling (incidence density sampling): controls are sampled at the exact time a case becomes a case, matched on person-time, making OR approximate the IRR.

Measure of effect

Odds ratio (OR); approximates RR when the disease is rare
  • Firestein & Kelley's Rheumatology, pp. 610-612

4. Cross-Sectional Studies

Measure exposure and outcome simultaneously in a defined population at a single point in time. Widely used for:
  • Drug utilization studies
  • Prevalence estimates of adverse drug effects
  • National surveys (e.g., NHANES, NHIS)

Limitations

  • Cannot establish temporal sequence (exposure before outcome), so causal inference is precluded.
  • Susceptible to prevalent user bias - people who died or discontinued treatment are already gone from the study population.

5. Special Designs in Pharmacoepidemiology

5a. Ecological Studies

Comparisons are made at the group or population level, not the individual level (e.g., comparing county-level medication sales with mortality rates). Useful for hypothesis generation. The major pitfall is the ecological fallacy - a group-level association may not hold at the individual level.
  • Rheumatology, 2-Volume Set (Elsevier), p. 229

5b. Case-Crossover Design

Each individual serves as their own control. For each case, a prior "control" time window from the same individual is selected and their exposure at that time is compared with exposure at the time of the event. This design:
  • Eliminates between-person confounding by fixed characteristics (genetics, sex, baseline risk)
  • Is particularly useful for transient exposures with acute outcomes (e.g., NSAIDs and MI, emotional triggers and ACS)
  • Requires that the exposure be intermittent and that carry-over effects are absent
To remove secular time trends, a case-time-control design runs a parallel case-crossover analysis in a control group.
  • Firestein & Kelley's Rheumatology, p. 611; Rheumatology, 2-Vol Set, p. 232

5c. New User Design (Incident User Design)

A methodological strategy - not a separate design per se - that restricts entry into a study to new initiators of a drug, excluding prevalent users. This:
  • Avoids depletion-of-susceptibles bias (patients who could not tolerate the drug have already quit)
  • Allows proper measurement of time-zero covariates before treatment starts
  • Provides a cleaner baseline from which to measure outcomes
  • Rheumatology, 2-Volume Set (Elsevier), p. 231

5d. Active Comparator Design

Instead of comparing a drug to unexposed (placebo-like) controls, the study compares two active drug users (e.g., statin A vs. statin B). This:
  • Controls for confounding by indication (indication for treatment is similar between groups)
  • Reduces channeling bias
  • Mirrors a head-to-head trial

6. Randomized Controlled Trials (RCTs) in Pharmacoepidemiology

The gold standard for establishing causation. Randomization distributes both measured and unmeasured confounders equally between groups. However, RCTs have major limitations in pharmacoepidemiology:
  • Too short for detecting long-term adverse effects
  • Exclude special populations (elderly, pregnant women, multi-morbid patients)
  • Insufficient power for rare adverse events
  • Pragmatic RCTs (conducted in routine care settings) bridge some of these gaps

7. Data Sources

The validity of any pharmacoepidemiological study depends heavily on the data source:
SourceExamplesStrengthsLimitations
Administrative claims databases (ACDs)Medicaid, Medicare, IBM MarketScan, OptumLarge, longitudinal, real-worldNo clinical variables (BP, BMI, labs); case ascertainment requires algorithms
Electronic health records (EHR)Hospital/clinic systemsRich multidimensional data (labs, notes, vitals, medications)Incomplete capture of out-of-network care
Disease registriesBiologics registries (e.g., BSRBR)Standardized data, focused populationMay lack comparator group
Vital statisticsBirth/death recordsComplete mortality dataLimited clinical detail
SurveysNHANES, NHISPopulation representativeCross-sectional; no causal inference
Wearables / digital biomarkersSmartwatch dataContinuous real-world monitoringValidation challenges
  • Firestein & Kelley's Rheumatology, pp. 608-609

8. Biases and Methodological Issues

Confounding

A third variable (confounder) is associated with both exposure and outcome, distorting the apparent association.
Confounding by indication is the most important bias in pharmacoepidemiology. Patients who receive an aggressive treatment often have more severe disease - so the drug appears harmful simply because sicker patients receive it. Control strategies:
  • Restriction (limit the study to one level of the confounder)
  • Matching
  • Multivariable regression
  • Propensity score methods (PS matching, PS stratification, inverse probability weighting) - balance measured confounders but cannot balance unmeasured ones
  • Firestein & Kelley's Rheumatology, p. 612 (confounding by indication)

Immortal Time Bias

Occurs when time during which a participant cannot experience the outcome is misclassified as exposed person-time. For example, if exposure status is only determined after a certain delay (e.g., requiring a subject to fill 3 prescriptions), the pre-classification period is "immortal" - yet it may be wrongly attributed to the exposed group, making the exposed group appear to have better outcomes. Corrected by using person-time analysis with proper time-zero definition.
  • Rheumatology, 2-Volume Set (Elsevier), p. 233

Selection Bias

Arises from how participants enter or remain in the study, causing the study sample to not represent the population of interest. Includes:
  • Healthy user / healthy adherer bias: Drug users who remain in studies may be healthier at baseline
  • Depletion of susceptibles: Early users who were most susceptible to adverse effects have already been removed from the study population
  • Berkson's bias: Hospital-based controls may over-represent the association between exposure and outcome

Misclassification

Any variable (exposure, outcome, confounder) may be incorrectly coded.
  • Non-differential (random) misclassification of a binary exposure biases the association toward the null (underestimation).
  • Differential misclassification can bias in either direction and is unpredictable.
  • Recall bias (a form of differential misclassification) is common in case-control studies using self-reported exposures.
  • Rheumatology, 2-Volume Set (Elsevier), p. 234

9. Measures of Association

DesignTypical Measure
CohortRisk ratio, Rate ratio (IRR), Hazard ratio
Case-controlOdds ratio (approximates RR when outcome is rare)
Cross-sectionalPrevalence ratio
AnyRisk difference, Attributable risk, Population attributable fraction
The standardized mortality ratio (SMR) compares observed deaths in a study population to those expected from a reference population (used in occupational pharmacoepidemiology).

10. Summary Table

DesignSelection byDirectionKey MeasureBest forMain Bias
Prospective cohortExposureForwardRR, HRCommon outcomes, multiple outcomesLoss to follow-up
Retrospective cohortExposure (past)Forward (using records)RR, IRROccupational exposuresInformation bias
Case-controlOutcomeBackwardORRare outcomesRecall bias, selection bias
Nested case-controlOutcome (within cohort)BackwardOR (≈IRR)Biomarkers, expensive assaysReduced
Case-crossoverOutcome onlySelf-comparisonORTransient exposures, acute outcomesTime-trend confounding
EcologicalGroup-levelEitherCorrelationHypothesis generationEcological fallacy
Cross-sectionalPoint in timeNonePRPrevalence, drug utilizationTemporal ambiguity

Key textbook sources used:
  • Firestein & Kelley's Textbook of Rheumatology (2-Volume Set) - Chapter 31 (Epidemiology)
  • Rheumatology, 2-Volume Set (Elsevier, 2022) - Chapter 23
  • Barash's Clinical Anesthesia, 9th Edition - Chapter 7 (Research Design)
This is a shared conversation. Sign in to Orris to start your own chat.