Designing for external validity
Slide 1The methods in the last two videos address external validity after the fact: given a completed study, can we transport its findings to a new population? But there's a prior question worth asking. How do we generate evidence that's transportable in the first place?
Transportability is a design problem, not only an analytical one
It begins before the first participant is enrolled
Choices about who participates, what gets measured, and what data exist on target populations
Those choices determine whether transportability analysis is even possible later
Slide 2Transportability isn't just an analytical problem to solve after data collection ends. It's a design problem that begins before the first participant is enrolled. The choices researchers make — about who participates, about what gets measured, and about what data exist on target populations — determine whether transportability analysis is even possible later.
Transportability analysis makes a simple but demanding requirement
You must know how effect modifiers are distributed in both populations
Which means measuring the right things in both places
Slide 3Transportability analysis requires knowing how effect modifiers are distributed in both the study population and the target population. That creates a simple but demanding requirement. You must measure the right things in both places.
Two ways that requirement fails
Suspect age modifies the effect but don't collect age in the trial — you can't assess it later
A ministry wants to know whether trial evidence applies, but no one has data on that population
Then there is nothing to transport to
Slide 4If you suspect that age modifies a treatment's effect but don't collect age data in your trial, you can't later assess whether age differences between populations explain divergent results. And if a ministry of health wants to know whether trial evidence applies to their population, but no one has collected data on that population's characteristics, there's nothing to transport to.
Three questions to ask before launching a study
What characteristics might plausibly modify how this treatment works?
Will comparable measurements exist in the populations we hope to inform?
Slide 5This is a design imperative rather than an afterthought. Before launching a study, researchers should ask: What characteristics might plausibly modify how this treatment works? Are we measuring them? Will comparable measurements exist in the populations we hope to inform? The answers won't always be clear — effect modification is often discovered rather than predicted. But the question itself focuses attention on external validity from the start.
Clinical trials have historically enrolled narrow populations
Women were excluded from many trials until the 1990s
Older adults are routinely excluded despite being the primary users of many treatments
Racial and ethnic minorities remain underrepresented across therapeutic areas
Slide 6Clinical trials have historically enrolled narrow populations. Women were excluded from many trials until the nineteen nineties, partly out of concerns about pregnancy, but also on the assumption that findings in men would apply to women. Older adults are routinely excluded despite being the primary users of many treatments. Racial and ethnic minorities remain underrepresented across therapeutic areas.
These exclusions are an equity concern, and also a scientific problem
Homogeneous trials cannot reveal effect heterogeneity
If everyone is young, male, and healthy, you can't learn whether age, sex, or comorbidities matter
You cannot identify effect modifiers you haven't studied
Slide 7These exclusions are often framed as equity concerns, and they are. But the problem is also scientific. Homogeneous trials cannot reveal effect heterogeneity. If everyone in your study is young, male, and healthy, you cannot learn whether age, sex, or comorbidities modify the treatment effect. You cannot identify effect modifiers you haven't studied, and you cannot assess transportability to populations you've excluded.
ACTG 320
The age analysis existed only because the trial enrolled across ages
Enroll only patients in their 30s, where treatment worked best
The overall estimate would have looked more impressive
And been far less informative about any other population
Slide 8Consider what this means for the HIV example. The researchers could examine whether treatment effects varied by age only because the trial had enrolled patients across age groups. Had it enrolled only patients in their thirties — the group where treatment worked best — the overall estimate would have looked even more impressive. It would also have been far less informative about what to expect in any other population.
Diversity in trials is the raw material for understanding heterogeneity
Not a box to check for regulatory approval
Inclusive enrollment produces evidence that reveals how treatments work across groups
Evidence that supports transportability rather than undermining it
Slide 9Diversity in trials isn't a box to check for regulatory approval. It's the raw material for understanding effect heterogeneity. Inclusive enrollment produces evidence that reveals how treatments work across groups — evidence that supports transportability rather than undermining it.
Real-world evidence describes the populations we want to reach
Electronic health records, insurance claims, disease registries, other routine sources
Its specific role in transportability is to characterize the target
Slide 10Real-world evidence — data from electronic health records, insurance claims, disease registries, and other routine sources — plays a specific role in transportability. It describes the target populations we want to reach.
Cole & Stuart, Am J Epidemiol 2010
The HIV analysis depended on data about the target
They needed the age, sex, and race distribution of the 2006 target population
CDC surveillance provided it
Without data on the target, the analysis would not have been possible
Slide 11When Cole and Stuart standardized the trial results to the later HIV population, they needed to know the age, sex, and race distribution of that target. CDC surveillance data provided it. Without data on the target population, the analysis wouldn't have been possible at all. The statistical methods are useless without something to transport to.
Data infrastructure is external validity infrastructure
Surveillance systems, registries, and linked administrative databases build the foundation
Systems without them face a harder problem
Evidence from elsewhere, and no way to formally assess whether it applies locally
Slide 12This is why data infrastructure matters for external validity. Countries and health systems that invest in routine data collection — surveillance systems, registries, linked administrative databases — create the foundation for transportability analysis. Those without such infrastructure face a harder problem. They may have evidence from elsewhere, but no way to formally assess whether it applies locally.
Real-world evidence also asks a different question
Trials establish efficacy
Real-world data reveal
Patients with multiple comorbidities
People who miss appointments
People who'd never enroll in a trial
Slide 13Real-world evidence serves another function as well. Trials establish efficacy under controlled conditions: selected patients, close monitoring, protocol-driven care. Real-world data reveal whether effects persist when treatments reach broader populations — patients with multiple comorbidities, those who miss appointments, people who would never have enrolled in a trial. This isn't replacing randomization with observation. It's asking a different question: Does the effect hold outside the conditions that produced it?
Regulators increasingly recognize this
The US FDA and the European Medicines Agency incorporate real-world evidence into certain decisions
Particularly for how treatments perform in populations underrepresented in trials
The goal is to raise the bar for external validity, not to lower it for causal inference
Slide 14Regulatory agencies increasingly recognize this. The US FDA and the European Medicines Agency now incorporate real-world evidence into certain decisions, particularly for understanding how treatments perform in populations underrepresented in trials. The goal isn't to lower the bar for causal inference. It's to raise the bar for external validity.
Shadish, Cook & Campbell, 2002
Five sources of external validity threat, framed as questions
Each is a way a causal relationship might fail to hold across variations in study conditions
Weighting and standardization treat problems that have already occurred
As in medicine, prevention is often better than treatment
Slide 15The methods we've discussed — weighting, standardization, transportability analysis — are treatments for external validity problems that have already occurred. But as in medicine, prevention is often better than treatment. The time to think about external validity is before the first participant is enrolled. A classic framework identifies five sources of external validity threat, each a way that causal relationships might fail to hold across variations in study conditions. Framed as questions, they become a design-stage checklist.
Threat 1 of 5 — interaction with units
Would the effect hold for different people?
Trials routinely exclude patients with comorbidities, limited literacy, or unstable housing
An intervention may work among healthy, motivated volunteers and fail among sicker patients
Ask: who are we excluding, and are those exclusions scientifically necessary or merely convenient?
Slide 16The first threat asks whether findings would differ if different kinds of participants had been studied. Clinical trials routinely exclude patients with comorbidities, limited literacy, or unstable living situations. An intervention may work beautifully among healthy, motivated volunteers and fail among sicker patients who face competing demands. In the HIV example, treatment effects varied substantially by age — information available only because the trial enrolled across age groups. So at the design stage, ask: Who are we excluding, and why? Are those exclusions scientifically necessary, or merely convenient?
Threat 2 of 5 — treatment variations
Would the effect hold for variations in the treatment?
Works with well-trained staff, supervised weekly and paid reliably
Fails at scale with minimal training, thin supervision, delayed payment
Ask: what are the active ingredients, and what might dilute them?
Slide 17The second threat asks whether findings depend on specific features of how the intervention was delivered. A community health worker program might work when health workers are well-trained, supervised weekly, and paid reliably. It may fail when implemented at scale with minimal training, infrequent supervision, and delayed payment. So at the design stage, ask: Are we testing a specific, replicable protocol, or an idealized version that won't survive contact with real health systems? What are the active ingredients, and what might dilute them?
Threat 3 of 5 — interaction with outcomes
Would the effect hold for a different outcome?
A training program might improve knowledge scores on a written test but not behavior change
An intervention might improve a biomarker but not the clinical endpoint patients care about
Ask: are we measuring what ultimately matters, or a convenient proxy?
Slide 18The third threat asks whether findings depend on how the outcome was measured. A training program might improve knowledge scores on a written test without changing actual behavior. A treatment might reduce symptoms as measured by clinician rating but not by patient self-report. An intervention might improve a biomarker but not the clinical endpoint patients actually care about. So at the design stage, ask: Are we measuring what ultimately matters, or a convenient proxy? Would stakeholders accept this outcome as meaningful?
Threat 4 of 5 — interaction with settings
Would the effect hold in a different setting?
An intervention tested in a well-resourced academic center may not work in an understaffed rural clinic
A finding from a system with universal coverage may not apply where patients pay out of pocket
Ask: what would need to be true about a setting for our findings to apply there?
Slide 19The fourth threat asks whether findings depend on features of where the study was conducted. An intervention tested in a well-resourced academic medical center may not work in an understaffed rural clinic. A finding from a health system with universal insurance coverage may not apply where patients pay out of pocket. So at the design stage, ask: What features of our study setting might not exist elsewhere? What would need to be true about a setting for our findings to apply there?
Threat 5 of 5 — context-dependent mediation
Would the same mechanism operate elsewhere?
In one setting, counseling raises uptake by improving knowledge
In another, by reducing stigma around seeking services
Ask: do we understand why this intervention should work, and are we measuring mediators?
Slide 20The fifth and final threat asks whether the reason an intervention works might differ across contexts, even when the overall effect looks similar. A family planning counseling program might increase contraceptive uptake in one setting by improving knowledge, and in another by reducing stigma around seeking services. If the mechanism differs, then information materials may matter less than community engagement. The same intervention may require different implementation strategies in different contexts. So at the design stage, ask: Do we understand why this intervention should work? Are we measuring mediators that would help us understand variation in effects across contexts?
Cole & Stuart, Am J Epidemiol 2010
Same drug, same outcome, different answer
Among trial participants
49% lower hazard of AIDS or death
In the target population
43% lower hazard of AIDS or death
Slide 21Return to Clinical Trial 320. A treatment that reduced the hazard of AIDS or death by forty-nine percent among trial participants would have reduced it by forty-three percent in the target population. Same drug, same outcome, different answer.
That gap isn't a failure of the trial
The researchers got the right answer for the people they studied
Causal effects have scope: they hold for particular populations, under particular conditions
Extending them beyond that scope requires assumptions
Slide 22That gap isn't a failure of the trial. The researchers got the right answer for the people they studied. The gap reflects something more fundamental. Causal effects have scope. They hold for particular populations, under particular conditions. Extending them beyond that scope requires assumptions — about which characteristics modify the effect, about how populations differ, about what stayed constant and what changed.
The framework, in four parts
Effect modification: only differences that change how treatments work
Generalizability and transportability: two versions of the problem
Statistical methods formalize assumptions; they don't eliminate them
The design-stage checklist comes before enrollment
Slide 23This chapter has given you a framework for reasoning about that scope. Effect modification is the key insight: population differences only matter when they involve characteristics that change how treatments work. Generalizability and transportability name two versions of the problem, depending on whether your sample sits inside the target population or outside it. Statistical methods can adjust for differences when you've measured the right variables, but they formalize assumptions rather than eliminate them. And the design-stage checklist belongs before the first participant is enrolled.
In global health this is the default condition
Trials from high-income countries inform guidelines implemented in low-resource settings
Studies from urban academic centers shape policy for rural health posts far away
Transport evidence without attention to effect modification, and the most different populations lose
Slide 24In global health, these concerns are not edge cases. They are the default condition. Evidence is routinely generated in one place and applied in another. Trials from high-income countries inform guidelines implemented in low-resource settings. Studies from urban academic centers shape policy for rural health posts thousands of kilometers away. When evidence is transported without attention to effect modification, the populations most different from trial participants may receive interventions that work less well for them, or don't work at all.
In Closing
Ask better questions
For whom was this shown? How might my population differ?
What would have to be true for this evidence to apply here?
Evidence is born in specific places, but decisions must live everywhere
Slide 25The goal isn't to become paralyzed by uncertainty. It's to ask better questions. For whom was this shown? How might my population differ? What would have to be true for this evidence to apply here? Evidence is born in specific places, but decisions must live everywhere. Your job is to reason carefully about the path between the two.