Designing for external validity Chapter 8: External Validity, Generalizability, and Transportability — Video 4 https://ghrbook.com/videos/designing-for-external-validity/ [Slide 1] The methods in the last two videos address external validity after the fact: given a completed study, can we transport its findings to a new population? But there's a prior question worth asking. How do we generate evidence that's transportable in the first place? [Slide 2] Transportability isn't just an analytical problem to solve after data collection ends. It's a design problem that begins before the first participant is enrolled. The choices researchers make — about who participates, about what gets measured, and about what data exist on target populations — determine whether transportability analysis is even possible later. [Slide 3] Transportability analysis requires knowing how effect modifiers are distributed in both the study population and the target population. That creates a simple but demanding requirement. You must measure the right things in both places. [Slide 4] If you suspect that age modifies a treatment's effect but don't collect age data in your trial, you can't later assess whether age differences between populations explain divergent results. And if a ministry of health wants to know whether trial evidence applies to their population, but no one has collected data on that population's characteristics, there's nothing to transport to. [Slide 5] This is a design imperative rather than an afterthought. Before launching a study, researchers should ask: What characteristics might plausibly modify how this treatment works? Are we measuring them? Will comparable measurements exist in the populations we hope to inform? The answers won't always be clear — effect modification is often discovered rather than predicted. But the question itself focuses attention on external validity from the start. [Slide 6] Clinical trials have historically enrolled narrow populations. Women were excluded from many trials until the nineteen nineties, partly out of concerns about pregnancy, but also on the assumption that findings in men would apply to women. Older adults are routinely excluded despite being the primary users of many treatments. Racial and ethnic minorities remain underrepresented across therapeutic areas. [Slide 7] These exclusions are often framed as equity concerns, and they are. But the problem is also scientific. Homogeneous trials cannot reveal effect heterogeneity. If everyone in your study is young, male, and healthy, you cannot learn whether age, sex, or comorbidities modify the treatment effect. You cannot identify effect modifiers you haven't studied, and you cannot assess transportability to populations you've excluded. [Slide 8] Consider what this means for the HIV example. The researchers could examine whether treatment effects varied by age only because the trial had enrolled patients across age groups. Had it enrolled only patients in their thirties — the group where treatment worked best — the overall estimate would have looked even more impressive. It would also have been far less informative about what to expect in any other population. [Slide 9] Diversity in trials isn't a box to check for regulatory approval. It's the raw material for understanding effect heterogeneity. Inclusive enrollment produces evidence that reveals how treatments work across groups — evidence that supports transportability rather than undermining it. [Slide 10] Real-world evidence — data from electronic health records, insurance claims, disease registries, and other routine sources — plays a specific role in transportability. It describes the target populations we want to reach. [Slide 11] When Cole and Stuart standardized the trial results to the later HIV population, they needed to know the age, sex, and race distribution of that target. CDC surveillance data provided it. Without data on the target population, the analysis wouldn't have been possible at all. The statistical methods are useless without something to transport to. [Slide 12] This is why data infrastructure matters for external validity. Countries and health systems that invest in routine data collection — surveillance systems, registries, linked administrative databases — create the foundation for transportability analysis. Those without such infrastructure face a harder problem. They may have evidence from elsewhere, but no way to formally assess whether it applies locally. [Slide 13] Real-world evidence serves another function as well. Trials establish efficacy under controlled conditions: selected patients, close monitoring, protocol-driven care. Real-world data reveal whether effects persist when treatments reach broader populations — patients with multiple comorbidities, those who miss appointments, people who would never have enrolled in a trial. This isn't replacing randomization with observation. It's asking a different question: Does the effect hold outside the conditions that produced it? [Slide 14] Regulatory agencies increasingly recognize this. The US FDA and the European Medicines Agency now incorporate real-world evidence into certain decisions, particularly for understanding how treatments perform in populations underrepresented in trials. The goal isn't to lower the bar for causal inference. It's to raise the bar for external validity. [Slide 15] The methods we've discussed — weighting, standardization, transportability analysis — are treatments for external validity problems that have already occurred. But as in medicine, prevention is often better than treatment. The time to think about external validity is before the first participant is enrolled. A classic framework identifies five sources of external validity threat, each a way that causal relationships might fail to hold across variations in study conditions. Framed as questions, they become a design-stage checklist. [Slide 16] The first threat asks whether findings would differ if different kinds of participants had been studied. Clinical trials routinely exclude patients with comorbidities, limited literacy, or unstable living situations. An intervention may work beautifully among healthy, motivated volunteers and fail among sicker patients who face competing demands. In the HIV example, treatment effects varied substantially by age — information available only because the trial enrolled across age groups. So at the design stage, ask: Who are we excluding, and why? Are those exclusions scientifically necessary, or merely convenient? [Slide 17] The second threat asks whether findings depend on specific features of how the intervention was delivered. A community health worker program might work when health workers are well-trained, supervised weekly, and paid reliably. It may fail when implemented at scale with minimal training, infrequent supervision, and delayed payment. So at the design stage, ask: Are we testing a specific, replicable protocol, or an idealized version that won't survive contact with real health systems? What are the active ingredients, and what might dilute them? [Slide 18] The third threat asks whether findings depend on how the outcome was measured. A training program might improve knowledge scores on a written test without changing actual behavior. A treatment might reduce symptoms as measured by clinician rating but not by patient self-report. An intervention might improve a biomarker but not the clinical endpoint patients actually care about. So at the design stage, ask: Are we measuring what ultimately matters, or a convenient proxy? Would stakeholders accept this outcome as meaningful? [Slide 19] The fourth threat asks whether findings depend on features of where the study was conducted. An intervention tested in a well-resourced academic medical center may not work in an understaffed rural clinic. A finding from a health system with universal insurance coverage may not apply where patients pay out of pocket. So at the design stage, ask: What features of our study setting might not exist elsewhere? What would need to be true about a setting for our findings to apply there? [Slide 20] The fifth and final threat asks whether the reason an intervention works might differ across contexts, even when the overall effect looks similar. A family planning counseling program might increase contraceptive uptake in one setting by improving knowledge, and in another by reducing stigma around seeking services. If the mechanism differs, then information materials may matter less than community engagement. The same intervention may require different implementation strategies in different contexts. So at the design stage, ask: Do we understand why this intervention should work? Are we measuring mediators that would help us understand variation in effects across contexts? [Slide 21] Return to Clinical Trial 320. A treatment that reduced the hazard of AIDS or death by forty-nine percent among trial participants would have reduced it by forty-three percent in the target population. Same drug, same outcome, different answer. [Slide 22] That gap isn't a failure of the trial. The researchers got the right answer for the people they studied. The gap reflects something more fundamental. Causal effects have scope. They hold for particular populations, under particular conditions. Extending them beyond that scope requires assumptions — about which characteristics modify the effect, about how populations differ, about what stayed constant and what changed. [Slide 23] This chapter has given you a framework for reasoning about that scope. Effect modification is the key insight: population differences only matter when they involve characteristics that change how treatments work. Generalizability and transportability name two versions of the problem, depending on whether your sample sits inside the target population or outside it. Statistical methods can adjust for differences when you've measured the right variables, but they formalize assumptions rather than eliminate them. And the design-stage checklist belongs before the first participant is enrolled. [Slide 24] In global health, these concerns are not edge cases. They are the default condition. Evidence is routinely generated in one place and applied in another. Trials from high-income countries inform guidelines implemented in low-resource settings. Studies from urban academic centers shape policy for rural health posts thousands of kilometers away. When evidence is transported without attention to effect modification, the populations most different from trial participants may receive interventions that work less well for them, or don't work at all. [Slide 25] The goal isn't to become paralyzed by uncertainty. It's to ask better questions. For whom was this shown? How might my population differ? What would have to be true for this evidence to apply here? Evidence is born in specific places, but decisions must live everywhere. Your job is to reason carefully about the path between the two.