Chapter 8 · Video 3

Do the results apply to a different population entirely?

25 slides · Video page · All videos · Transcript
Print layout
Slide 1

Do the results apply to a different population entirely?

Slide 2

The sample sits outside the population we want to inform

Generalizability
Sample inside the target
The population it was drawn from
Video 2
Transportability
Sample outside the target
A population never sampled
This video
Slide 3

Three versions of the same question

A ministry of health in Kenya reads a trial conducted in India
A funder reviews evidence from urban clinics and asks about rural settings never included
A policymaker sees results from a research-intensive trial and asks about routine conditions
Slide 4

No improvement to the original sampling design helps here

The Kenyan ministry isn't asking whether the Indian trial represented its source population
They're asking whether findings from India apply in Kenya
The study population and the target population don't overlap at all
Slide 5

The same effect modification logic, on a harder problem

Generalizability asks
Same target the sample came from
Are modifiers distributed differently?
Transportability asks
Two entirely separate populations
Are modifiers distributed differently?
Would the mechanism operate the same?
Slide 6

Transportability requires reasoning about how an intervention works

Not only whether it worked
Understand the mechanism, and the conditions under which it operates
Then reason about whether those conditions are likely to hold elsewhere
Slide 7
Hammer et al., NEJM 1997 — ACTG 320

A randomized trial in the mid-1990s

1,156 adults living with HIV in the United States and Puerto Rico
All immunosuppressed at entry: median CD4 count of 75 cells per cubic millimeter
Combination antiretroviral therapy against the standard treatment of the time
AIDS or death fell from 11% to 6%; hazard ratio 0.50
Slide 8

The trial had strong internal validity

Randomization ensured that, among participants, the estimated treatment effect was unbiased
The question that emerged a decade later was about whether the result still applied
Slide 9

By the mid-2000s the population looked very different

Advances in testing meant many people were diagnosed earlier
Treatment guidelines had evolved
Patients were younger, more racially diverse, often less immunosuppressed when treatment began
Slide 10

Applying the trial estimate directly would assume too much

It assumes the treatment effect is the same for everyone, regardless of age or disease severity
For antiretroviral therapy, baseline immune status matters
Patients starting with very low counts may gain more than those treated earlier in the disease
Slide 11
Cole & Stuart, Am J Epidemiol 2010

They asked a counterfactual question instead

What would the trial have shown if it had been conducted in the later population?
The target: the estimated 54,220 people newly infected with HIV in the US in 2006
Recently infected people have relatively normal immune function
Slide 12
Cole & Stuart, Am J Epidemiol 2010

Answering it requires two ingredients

Identify effect modifiers — characteristics that change how well the treatment works
Know how those characteristics are distributed in both populations
They had the joint distribution of three: age, sex, and race
Slide 13
Cole & Stuart, Am J Epidemiol 2010

The two populations differed substantially

Trial (n = 1,156)
9% were under 30
54% white, 28% Black
83% male
Target (54,220)
34% were under 30
36% white, 46% Black
73% male
Slide 14
Cole & Stuart, Am J Epidemiol 2010
Did the treatment effect vary by age?
Hazard ratios for AIDS or death within one year, ACTG 320, overall and by age group, with the age composition of the trial and of the target population. P for homogeneity = 0.03.
Slide 15

Now the population differences matter

The trial was dominated by patients in their 30s and 40s — where treatment worked best
The target population was younger, with over a third under age 30
9% of the trial, 34% of the target
Slide 16

Reweighting re-expresses the effect under the target's conditions

Give more weight to trial participants who resemble people in the later population
Give less weight to those who do not
The goal isn't to fix the trial or make it representative in a sampling sense
Slide 17
Cole & Stuart, Am J Epidemiol 2010

The estimate moved, and moved toward the null

Trial (intent-to-treat)
Hazard ratio 0.51
= a 49% lower hazard
Weighted to the target
Hazard ratio 0.57
= a 43% lower hazard
Muted by about 12%
Slide 18

Lack of representativeness is not a flaw

ACTG 320 was never designed to represent people who would be infected a decade later
It did not need to be
Its internal validity was strong throughout
Slide 19

Differences between populations matter only when they modify the effect

Many characteristics differed between the trial and the target
Only those that changed how treatment worked were relevant for transportability
Slide 20

Transportability is not automatic

It requires explicit assumptions about which features of the population matter
And data on those features in both populations
Immune status was the modifier they could not measure in the target
Slide 21

These methods address one source of external validity bias

Differences in who receives treatment, when treatment effects are heterogeneous
They don't address changes in the treatment itself
Or different outcome definitions, or shifts in health system context
Slide 22

Four questions to ask before assuming evidence transports

On what characteristics do the populations differ?
Which of those might modify the treatment effect, and why?
How large are the differences on those effect modifiers?
Can we adjust — are the modifiers measured in both populations?
Slide 23

These questions won't always have clean answers

Often we don't know which characteristics modify effects
Or we can't measure them in both populations, or we're uncertain about mechanisms
Asking explicitly beats assuming evidence transports unchanged
Slide 24

In global health, transportability is the norm

Evidence is routinely generated in one setting and applied in another
Across countries, health systems, and historical periods
The work is making explicit the assumptions required to extend a causal claim
Slide 25
In Closing

The logic holds even when formal adjustment is out of reach

Identify what differs, and reason about whether those differences modify effects
Then be honest about what you don't know

Do the results apply to a different population entirely?

Slide 1Transportability is the harder of the two external validity problems. It asks whether findings apply to a population the study never sampled from — where the study population sits outside the target population entirely.

The sample sits outside the population we want to inform

Generalizability
Sample inside the target
The population it was drawn from
Video 2
Transportability
Sample outside the target
A population never sampled
This video
Slide 2Generalizability, from the last video, addresses cases where the study sample is contained within the target population. Transportability addresses the case where it isn't — where the results are being applied to a population the study never touched.

Three versions of the same question

A ministry of health in Kenya reads a trial conducted in India
A funder reviews evidence from urban clinics and asks about rural settings never included
A policymaker sees results from a research-intensive trial and asks about routine conditions
Slide 3A ministry of health in Kenya reads a trial conducted in India and wonders whether the intervention would work in their context. A funder reviews evidence from urban clinics and asks whether it applies to rural settings that were never included. A policymaker sees results from a research-intensive trial and asks whether they would hold under routine conditions in a different health system.

No improvement to the original sampling design helps here

The Kenyan ministry isn't asking whether the Indian trial represented its source population
They're asking whether findings from India apply in Kenya
The study population and the target population don't overlap at all
Slide 4No improvement to the original study's sampling design answers these questions. The Kenyan ministry isn't asking whether the Indian trial represented its source population. They're asking whether findings from India apply in Kenya. The study population and the target population don't overlap at all.

The same effect modification logic, on a harder problem

Generalizability asks
Same target the sample came from
Are modifiers distributed differently?
Transportability asks
Two entirely separate populations
Are modifiers distributed differently?
Would the mechanism operate the same?
Slide 5The same effect modification logic applies, but the problem is more challenging. For generalizability, you're asking whether effect modifiers are distributed differently in your sample than in the same target population it was drawn from. For transportability, you're asking whether effect modifiers are distributed differently across entirely separate populations — and whether the causal mechanism that produced the effect in one setting would operate the same way in another.

Transportability requires reasoning about how an intervention works

Not only whether it worked
Understand the mechanism, and the conditions under which it operates
Then reason about whether those conditions are likely to hold elsewhere
Slide 6This is why transportability requires reasoning about how an intervention works, not just whether it worked. If you understand the mechanism, and the conditions under which it operates, you can reason about whether those conditions are likely to hold elsewhere.
Hammer et al., NEJM 1997 — ACTG 320

A randomized trial in the mid-1990s

1,156 adults living with HIV in the United States and Puerto Rico
All immunosuppressed at entry: median CD4 count of 75 cells per cubic millimeter
Combination antiretroviral therapy against the standard treatment of the time
AIDS or death fell from 11% to 6%; hazard ratio 0.50
Slide 7To see how transportability works in practice, take the HIV example from the first video and open it up. In the mid nineteen nineties, a randomized clinical trial known as Clinical Trial 320 evaluated combination antiretroviral therapy among adults living with HIV in the United States. It enrolled one thousand one hundred fifty-six patients, all severely immunosuppressed at entry, and combination therapy cut the risk of AIDS or death roughly in half. This was one of the studies that established combination therapy as standard care.

The trial had strong internal validity

Randomization ensured that, among participants, the estimated treatment effect was unbiased
The question that emerged a decade later was about whether the result still applied
Slide 8The trial had strong internal validity. Randomization ensured that, among participants, the estimated effect of treatment was unbiased. But the question that emerged a decade later was not about whether the trial was correct. It was about whether the result still applied.

By the mid-2000s the population looked very different

Advances in testing meant many people were diagnosed earlier
Treatment guidelines had evolved
Patients were younger, more racially diverse, often less immunosuppressed when treatment began
Slide 9By the mid-two-thousands, the population of people living with HIV in the United States looked very different from the population enrolled in the trial. Advances in testing meant that many people were diagnosed earlier. Treatment guidelines had evolved. Patients were younger, more racially diverse, and often less immunosuppressed at the time treatment began. Public health officials wanted to know: would the dramatic benefits observed in the trial still hold for the broader HIV population a decade later?

Applying the trial estimate directly would assume too much

It assumes the treatment effect is the same for everyone, regardless of age or disease severity
For antiretroviral therapy, baseline immune status matters
Patients starting with very low counts may gain more than those treated earlier in the disease
Slide 10One tempting response would be to simply apply the trial's effect estimate to the contemporary HIV population. But that assumes the treatment effect is the same for everyone, regardless of age, disease severity, or the other characteristics that changed over time. That assumption is unlikely to hold. For antiretroviral therapy, baseline immune status matters: patients with very low immune cell counts may gain more from treatment than those who begin earlier in the disease course. The two populations differed not just in who they included, but in characteristics that plausibly modify the treatment effect.
Cole & Stuart, Am J Epidemiol 2010

They asked a counterfactual question instead

What would the trial have shown if it had been conducted in the later population?
The target: the estimated 54,220 people newly infected with HIV in the US in 2006
Recently infected people have relatively normal immune function
Slide 11Cole and Stuart approached this by asking a counterfactual question. What would the trial have shown if it had been conducted in the later population instead of the original trial population? Their target population was specific: the roughly fifty-four thousand people estimated to have been newly infected with HIV in the United States in two thousand six. That specificity matters for everything that follows. Recently infected people have relatively normal immune function, and the trial sample was severely immunosuppressed.
Cole & Stuart, Am J Epidemiol 2010

Answering it requires two ingredients

Identify effect modifiers — characteristics that change how well the treatment works
Know how those characteristics are distributed in both populations
They had the joint distribution of three: age, sex, and race
Slide 12Answering that question requires two ingredients. First, identify effect modifiers — characteristics that change how well the treatment works. Second, know how those characteristics are distributed in both populations. For the trial they had individual patient records. For the target population they had no individual-level data, only the joint distribution of three characteristics: age, sex, and race. Those three are what the analysis could use.
Cole & Stuart, Am J Epidemiol 2010

The two populations differed substantially

Trial (n = 1,156)
9% were under 30
54% white, 28% Black
83% male
Target (54,220)
34% were under 30
36% white, 46% Black
73% male
Slide 13And the populations differed substantially. The trial over-sampled older patients: ninety-one percent were age thirty or older, compared with sixty-six percent in the target population. It over-sampled white patients, fifty-four percent against thirty-six, and under-sampled Black patients, twenty-eight percent against forty-six. If the treatment effect differs across levels of these characteristics, and the distributions differ between populations, then the original estimate will not apply directly.
Cole & Stuart, Am J Epidemiol 2010
Did the treatment effect vary by age?
Hazard ratios for AIDS or death within one year, ACTG 320, overall and by age group, with the age composition of the trial and of the target population. P for homogeneity = 0.03.
Slide 14So did the treatment effect vary? When the researchers examined age-stratified results, they found striking heterogeneity. Among patients in their thirties, treatment was dramatically protective — a hazard ratio of zero point two one, roughly an eighty percent reduction. Among patients in their forties the effect was much weaker, and among the youngest patients treatment appeared to increase risk, though that estimate was imprecise given the small subgroup. You can see that imprecision in the interval: it runs from zero point three four to ten point two. The overall estimate of zero point five one masked all of this variation.

Now the population differences matter

The trial was dominated by patients in their 30s and 40s — where treatment worked best
The target population was younger, with over a third under age 30
9% of the trial, 34% of the target
Slide 15Now the population differences matter. The trial was dominated by patients in their thirties and forties, the groups where treatment worked best. The target population was younger, with over a third under age thirty. The age band that made up nine percent of the trial made up thirty-four percent of the population the result was headed for.

Reweighting re-expresses the effect under the target's conditions

Give more weight to trial participants who resemble people in the later population
Give less weight to those who do not
The goal isn't to fix the trial or make it representative in a sampling sense
Slide 16So rather than discarding the trial results, Cole and Stuart showed how the trial data could be reweighted, so that the distribution of key effect modifiers matched that of the target population. In practice this means giving more weight to trial participants who resemble people in the later HIV population, and less weight to those who do not. The goal wasn't to fix the trial or to make it representative in a sampling sense. The goal was to re-express the causal effect under the conditions that define the target population.
Cole & Stuart, Am J Epidemiol 2010

The estimate moved, and moved toward the null

Trial (intent-to-treat)
Hazard ratio 0.51
= a 49% lower hazard
Weighted to the target
Hazard ratio 0.57
= a 43% lower hazard
Muted by about 12%
Slide 17When the adjustment was applied, the estimate changed. The trial's intent-to-treat hazard ratio was zero point five one. Weighted to the target population, it was zero point five seven — moved closer to one, which means a shrinking effect. The authors describe it as muted by about twelve percent. And these are the two numbers the first video showed as percentages: a hazard ratio of zero point five one is a forty-nine percent lower hazard of AIDS or death, and zero point five seven is forty-three percent. Same analysis, same paper, expressed on a different scale. The benefit of combination therapy was still substantial, but attenuated — not because the trial was biased, but because the population it was being applied to had changed in ways that mattered.

Lack of representativeness is not a flaw

ACTG 320 was never designed to represent people who would be infected a decade later
It did not need to be
Its internal validity was strong throughout
Slide 18The example teaches several things at once. The first: a lack of representativeness is not a flaw. The trial was never designed to represent all people living with HIV a decade in the future, and it did not need to be. Its internal validity was strong.

Differences between populations matter only when they modify the effect

Many characteristics differed between the trial and the target
Only those that changed how treatment worked were relevant for transportability
Slide 19The second: differences between populations only matter when they modify the effect. Many characteristics differed between the trial and target populations, but only the ones that changed how well treatment worked were relevant here.

Transportability is not automatic

It requires explicit assumptions about which features of the population matter
And data on those features in both populations
Immune status was the modifier they could not measure in the target
Slide 20The third: transportability is not automatic. It requires explicit assumptions about which features of the population matter, and data on those features in both populations. Here the gap is a real one. Baseline immune status is clinically the most plausible modifier of this treatment's effect, and it is exactly the characteristic they had no data on for the target population. The authors say so directly. Transportability analysis is limited there, not by a statistical failure, but by missing information.

These methods address one source of external validity bias

Differences in who receives treatment, when treatment effects are heterogeneous
They don't address changes in the treatment itself
Or different outcome definitions, or shifts in health system context
Slide 21The fourth: these methods address only one source of external validity bias — differences in who receives treatment, when treatment effects are heterogeneous. They do not address changes in the treatment itself, differences in how outcomes are defined, or shifts in health system context. Those require substantive judgment rather than statistical adjustment.

Four questions to ask before assuming evidence transports

On what characteristics do the populations differ?
Which of those might modify the treatment effect, and why?
How large are the differences on those effect modifiers?
Can we adjust — are the modifiers measured in both populations?
Slide 22The example generalizes into four questions worth asking before assuming evidence from one population applies to another. On what characteristics do the populations differ? Which of those might modify the treatment effect, and why? How large are the differences on those modifiers? And can we adjust — are they measured in both populations?

These questions won't always have clean answers

Often we don't know which characteristics modify effects
Or we can't measure them in both populations, or we're uncertain about mechanisms
Asking explicitly beats assuming evidence transports unchanged
Slide 23These questions won't always have clean answers. Often we don't know which characteristics modify effects, or we can't measure them in both populations, or we're uncertain about mechanisms. But asking them explicitly is better than assuming evidence transports unchanged, or assuming it doesn't transport at all.

In global health, transportability is the norm

Evidence is routinely generated in one setting and applied in another
Across countries, health systems, and historical periods
The work is making explicit the assumptions required to extend a causal claim
Slide 24In global health, transportability is not an edge case. It is the norm. Evidence is routinely generated in one setting and applied in another: across countries, across health systems, and across historical periods. The work is making explicit the assumptions required to extend a causal claim from one population to another, and recognizing when those assumptions are plausible, questionable, or untenable.
In Closing

The logic holds even when formal adjustment is out of reach

Identify what differs, and reason about whether those differences modify effects
Then be honest about what you don't know
Slide 25Transportability methods offer one way to formalize this reasoning when the necessary data are available. But even when formal adjustment isn't possible, the underlying logic still holds. Identify what differs. Reason about whether those differences modify effects. And be honest about what you don't know.