Do the results apply to a different population entirely? Chapter 8: External Validity, Generalizability, and Transportability — Video 3 https://ghrbook.com/videos/do-results-apply-to-another-population/ [Slide 1] Transportability is the harder of the two external validity problems. It asks whether findings apply to a population the study never sampled from — where the study population sits outside the target population entirely. [Slide 2] Generalizability, from the last video, addresses cases where the study sample is contained within the target population. Transportability addresses the case where it isn't — where the results are being applied to a population the study never touched. [Slide 3] A ministry of health in Kenya reads a trial conducted in India and wonders whether the intervention would work in their context. A funder reviews evidence from urban clinics and asks whether it applies to rural settings that were never included. A policymaker sees results from a research-intensive trial and asks whether they would hold under routine conditions in a different health system. [Slide 4] No improvement to the original study's sampling design answers these questions. The Kenyan ministry isn't asking whether the Indian trial represented its source population. They're asking whether findings from India apply in Kenya. The study population and the target population don't overlap at all. [Slide 5] The same effect modification logic applies, but the problem is more challenging. For generalizability, you're asking whether effect modifiers are distributed differently in your sample than in the same target population it was drawn from. For transportability, you're asking whether effect modifiers are distributed differently across entirely separate populations — and whether the causal mechanism that produced the effect in one setting would operate the same way in another. [Slide 6] This is why transportability requires reasoning about how an intervention works, not just whether it worked. If you understand the mechanism, and the conditions under which it operates, you can reason about whether those conditions are likely to hold elsewhere. [Slide 7] To see how transportability works in practice, take the HIV example from the first video and open it up. In the mid nineteen nineties, a randomized clinical trial known as Clinical Trial 320 evaluated combination antiretroviral therapy among adults living with HIV in the United States. It enrolled one thousand one hundred fifty-six patients, all severely immunosuppressed at entry, and combination therapy cut the risk of AIDS or death roughly in half. This was one of the studies that established combination therapy as standard care. [Slide 8] The trial had strong internal validity. Randomization ensured that, among participants, the estimated effect of treatment was unbiased. But the question that emerged a decade later was not about whether the trial was correct. It was about whether the result still applied. [Slide 9] By the mid-two-thousands, the population of people living with HIV in the United States looked very different from the population enrolled in the trial. Advances in testing meant that many people were diagnosed earlier. Treatment guidelines had evolved. Patients were younger, more racially diverse, and often less immunosuppressed at the time treatment began. Public health officials wanted to know: would the dramatic benefits observed in the trial still hold for the broader HIV population a decade later? [Slide 10] One tempting response would be to simply apply the trial's effect estimate to the contemporary HIV population. But that assumes the treatment effect is the same for everyone, regardless of age, disease severity, or the other characteristics that changed over time. That assumption is unlikely to hold. For antiretroviral therapy, baseline immune status matters: patients with very low immune cell counts may gain more from treatment than those who begin earlier in the disease course. The two populations differed not just in who they included, but in characteristics that plausibly modify the treatment effect. [Slide 11] Cole and Stuart approached this by asking a counterfactual question. What would the trial have shown if it had been conducted in the later population instead of the original trial population? Their target population was specific: the roughly fifty-four thousand people estimated to have been newly infected with HIV in the United States in two thousand six. That specificity matters for everything that follows. Recently infected people have relatively normal immune function, and the trial sample was severely immunosuppressed. [Slide 12] Answering that question requires two ingredients. First, identify effect modifiers — characteristics that change how well the treatment works. Second, know how those characteristics are distributed in both populations. For the trial they had individual patient records. For the target population they had no individual-level data, only the joint distribution of three characteristics: age, sex, and race. Those three are what the analysis could use. [Slide 13] And the populations differed substantially. The trial over-sampled older patients: ninety-one percent were age thirty or older, compared with sixty-six percent in the target population. It over-sampled white patients, fifty-four percent against thirty-six, and under-sampled Black patients, twenty-eight percent against forty-six. If the treatment effect differs across levels of these characteristics, and the distributions differ between populations, then the original estimate will not apply directly. [Slide 14] So did the treatment effect vary? When the researchers examined age-stratified results, they found striking heterogeneity. Among patients in their thirties, treatment was dramatically protective — a hazard ratio of zero point two one, roughly an eighty percent reduction. Among patients in their forties the effect was much weaker, and among the youngest patients treatment appeared to increase risk, though that estimate was imprecise given the small subgroup. You can see that imprecision in the interval: it runs from zero point three four to ten point two. The overall estimate of zero point five one masked all of this variation. [Slide 15] Now the population differences matter. The trial was dominated by patients in their thirties and forties, the groups where treatment worked best. The target population was younger, with over a third under age thirty. The age band that made up nine percent of the trial made up thirty-four percent of the population the result was headed for. [Slide 16] So rather than discarding the trial results, Cole and Stuart showed how the trial data could be reweighted, so that the distribution of key effect modifiers matched that of the target population. In practice this means giving more weight to trial participants who resemble people in the later HIV population, and less weight to those who do not. The goal wasn't to fix the trial or to make it representative in a sampling sense. The goal was to re-express the causal effect under the conditions that define the target population. [Slide 17] When the adjustment was applied, the estimate changed. The trial's intent-to-treat hazard ratio was zero point five one. Weighted to the target population, it was zero point five seven — moved closer to one, which means a shrinking effect. The authors describe it as muted by about twelve percent. And these are the two numbers the first video showed as percentages: a hazard ratio of zero point five one is a forty-nine percent lower hazard of AIDS or death, and zero point five seven is forty-three percent. Same analysis, same paper, expressed on a different scale. The benefit of combination therapy was still substantial, but attenuated — not because the trial was biased, but because the population it was being applied to had changed in ways that mattered. [Slide 18] The example teaches several things at once. The first: a lack of representativeness is not a flaw. The trial was never designed to represent all people living with HIV a decade in the future, and it did not need to be. Its internal validity was strong. [Slide 19] The second: differences between populations only matter when they modify the effect. Many characteristics differed between the trial and target populations, but only the ones that changed how well treatment worked were relevant here. [Slide 20] The third: transportability is not automatic. It requires explicit assumptions about which features of the population matter, and data on those features in both populations. Here the gap is a real one. Baseline immune status is clinically the most plausible modifier of this treatment's effect, and it is exactly the characteristic they had no data on for the target population. The authors say so directly. Transportability analysis is limited there, not by a statistical failure, but by missing information. [Slide 21] The fourth: these methods address only one source of external validity bias — differences in who receives treatment, when treatment effects are heterogeneous. They do not address changes in the treatment itself, differences in how outcomes are defined, or shifts in health system context. Those require substantive judgment rather than statistical adjustment. [Slide 22] The example generalizes into four questions worth asking before assuming evidence from one population applies to another. On what characteristics do the populations differ? Which of those might modify the treatment effect, and why? How large are the differences on those modifiers? And can we adjust — are they measured in both populations? [Slide 23] These questions won't always have clean answers. Often we don't know which characteristics modify effects, or we can't measure them in both populations, or we're uncertain about mechanisms. But asking them explicitly is better than assuming evidence transports unchanged, or assuming it doesn't transport at all. [Slide 24] In global health, transportability is not an edge case. It is the norm. Evidence is routinely generated in one setting and applied in another: across countries, across health systems, and across historical periods. The work is making explicit the assumptions required to extend a causal claim from one population to another, and recognizing when those assumptions are plausible, questionable, or untenable. [Slide 25] Transportability methods offer one way to formalize this reasoning when the necessary data are available. But even when formal adjustment isn't possible, the underlying logic still holds. Identify what differs. Reason about whether those differences modify effects. And be honest about what you don't know.