Video 4 of 4
Designing for external validity
External validity as a design decision made before enrollment rather than an analysis run afterward, then Shadish, Cook and Campbell’s five threats as a checklist you can put to your own design or to someone else’s paper.
9:57 · 25 slides · printable slides · transcript
Slides
Printable deck →▶Transcript25 sections
Generated from the narration script. Plain text version.
1The methods in the last two videos address external validity after the fact: given a completed study, can we transport its findings to a new population? But there's a prior question worth asking. How do we generate evidence that's transportable in the first place?
2Transportability isn't just an analytical problem to solve after data collection ends. It's a design problem that begins before the first participant is enrolled. The choices researchers make — about who participates, about what gets measured, and about what data exist on target populations — determine whether transportability analysis is even possible later.
3Transportability analysis requires knowing how effect modifiers are distributed in both the study population and the target population. That creates a simple but demanding requirement. You must measure the right things in both places.
4If you suspect that age modifies a treatment's effect but don't collect age data in your trial, you can't later assess whether age differences between populations explain divergent results. And if a ministry of health wants to know whether trial evidence applies to their population, but no one has collected data on that population's characteristics, there's nothing to transport to.
5This is a design imperative rather than an afterthought. Before launching a study, researchers should ask: What characteristics might plausibly modify how this treatment works? Are we measuring them? Will comparable measurements exist in the populations we hope to inform? The answers won't always be clear — effect modification is often discovered rather than predicted. But the question itself focuses attention on external validity from the start.
6Clinical trials have historically enrolled narrow populations. Women were excluded from many trials until the nineteen nineties, partly out of concerns about pregnancy, but also on the assumption that findings in men would apply to women. Older adults are routinely excluded despite being the primary users of many treatments. Racial and ethnic minorities remain underrepresented across therapeutic areas.
7These exclusions are often framed as equity concerns, and they are. But the problem is also scientific. Homogeneous trials cannot reveal effect heterogeneity. If everyone in your study is young, male, and healthy, you cannot learn whether age, sex, or comorbidities modify the treatment effect. You cannot identify effect modifiers you haven't studied, and you cannot assess transportability to populations you've excluded.
8Consider what this means for the HIV example. The researchers could examine whether treatment effects varied by age only because the trial had enrolled patients across age groups. Had it enrolled only patients in their thirties — the group where treatment worked best — the overall estimate would have looked even more impressive. It would also have been far less informative about what to expect in any other population.
9Diversity in trials isn't a box to check for regulatory approval. It's the raw material for understanding effect heterogeneity. Inclusive enrollment produces evidence that reveals how treatments work across groups — evidence that supports transportability rather than undermining it.
10Real-world evidence — data from electronic health records, insurance claims, disease registries, and other routine sources — plays a specific role in transportability. It describes the target populations we want to reach.
11When Cole and Stuart standardized the trial results to the later HIV population, they needed to know the age, sex, and race distribution of that target. CDC surveillance data provided it. Without data on the target population, the analysis wouldn't have been possible at all. The statistical methods are useless without something to transport to.
12This is why data infrastructure matters for external validity. Countries and health systems that invest in routine data collection — surveillance systems, registries, linked administrative databases — create the foundation for transportability analysis. Those without such infrastructure face a harder problem. They may have evidence from elsewhere, but no way to formally assess whether it applies locally.
13Real-world evidence serves another function as well. Trials establish efficacy under controlled conditions: selected patients, close monitoring, protocol-driven care. Real-world data reveal whether effects persist when treatments reach broader populations — patients with multiple comorbidities, those who miss appointments, people who would never have enrolled in a trial. This isn't replacing randomization with observation. It's asking a different question: Does the effect hold outside the conditions that produced it?
14Regulatory agencies increasingly recognize this. The US FDA and the European Medicines Agency now incorporate real-world evidence into certain decisions, particularly for understanding how treatments perform in populations underrepresented in trials. The goal isn't to lower the bar for causal inference. It's to raise the bar for external validity.
15The methods we've discussed — weighting, standardization, transportability analysis — are treatments for external validity problems that have already occurred. But as in medicine, prevention is often better than treatment. The time to think about external validity is before the first participant is enrolled. A classic framework identifies five sources of external validity threat, each a way that causal relationships might fail to hold across variations in study conditions. Framed as questions, they become a design-stage checklist.
16The first threat asks whether findings would differ if different kinds of participants had been studied. Clinical trials routinely exclude patients with comorbidities, limited literacy, or unstable living situations. An intervention may work beautifully among healthy, motivated volunteers and fail among sicker patients who face competing demands. In the HIV example, treatment effects varied substantially by age — information available only because the trial enrolled across age groups. So at the design stage, ask: Who are we excluding, and why? Are those exclusions scientifically necessary, or merely convenient?
17The second threat asks whether findings depend on specific features of how the intervention was delivered. A community health worker program might work when health workers are well-trained, supervised weekly, and paid reliably. It may fail when implemented at scale with minimal training, infrequent supervision, and delayed payment. So at the design stage, ask: Are we testing a specific, replicable protocol, or an idealized version that won't survive contact with real health systems? What are the active ingredients, and what might dilute them?
18The third threat asks whether findings depend on how the outcome was measured. A training program might improve knowledge scores on a written test without changing actual behavior. A treatment might reduce symptoms as measured by clinician rating but not by patient self-report. An intervention might improve a biomarker but not the clinical endpoint patients actually care about. So at the design stage, ask: Are we measuring what ultimately matters, or a convenient proxy? Would stakeholders accept this outcome as meaningful?
19The fourth threat asks whether findings depend on features of where the study was conducted. An intervention tested in a well-resourced academic medical center may not work in an understaffed rural clinic. A finding from a health system with universal insurance coverage may not apply where patients pay out of pocket. So at the design stage, ask: What features of our study setting might not exist elsewhere? What would need to be true about a setting for our findings to apply there?
20The fifth and final threat asks whether the reason an intervention works might differ across contexts, even when the overall effect looks similar. A family planning counseling program might increase contraceptive uptake in one setting by improving knowledge, and in another by reducing stigma around seeking services. If the mechanism differs, then information materials may matter less than community engagement. The same intervention may require different implementation strategies in different contexts. So at the design stage, ask: Do we understand why this intervention should work? Are we measuring mediators that would help us understand variation in effects across contexts?
21Return to Clinical Trial 320. A treatment that reduced the hazard of AIDS or death by forty-nine percent among trial participants would have reduced it by forty-three percent in the target population. Same drug, same outcome, different answer.
22That gap isn't a failure of the trial. The researchers got the right answer for the people they studied. The gap reflects something more fundamental. Causal effects have scope. They hold for particular populations, under particular conditions. Extending them beyond that scope requires assumptions — about which characteristics modify the effect, about how populations differ, about what stayed constant and what changed.
23This chapter has given you a framework for reasoning about that scope. Effect modification is the key insight: population differences only matter when they involve characteristics that change how treatments work. Generalizability and transportability name two versions of the problem, depending on whether your sample sits inside the target population or outside it. Statistical methods can adjust for differences when you've measured the right variables, but they formalize assumptions rather than eliminate them. And the design-stage checklist belongs before the first participant is enrolled.
24In global health, these concerns are not edge cases. They are the default condition. Evidence is routinely generated in one place and applied in another. Trials from high-income countries inform guidelines implemented in low-resource settings. Studies from urban academic centers shape policy for rural health posts thousands of kilometers away. When evidence is transported without attention to effect modification, the populations most different from trial participants may receive interventions that work less well for them, or don't work at all.
25The goal isn't to become paralyzed by uncertainty. It's to ask better questions. For whom was this shown? How might my population differ? What would have to be true for this evidence to apply here? Evidence is born in specific places, but decisions must live everywhere. Your job is to reason carefully about the path between the two.