Does the sample stand for the population it came from?
Slide 1Generalizability is the first of the two external validity problems. It asks whether results from a study sample apply to the target population that sample was drawn from. Its defining characteristic is containment: everyone in the study sample is also a member of the target population.
The sample sits inside the population we want to inform
Generalizability
Everyone studied is in the target
A Kenyan trial, informing Kenya
Transportability
Sample outside the target
A different population entirely
A Kenyan trial, informing Uganda
Slide 2That containment is what separates this problem from the one in the next video. In generalizability, everyone we studied belongs to the population we want to inform. In transportability, they don't. Everything that follows here assumes the first arrangement.
One way to achieve generalizability is representative sampling
Draw the sample randomly from the target, or by probability methods with known selection probabilities
On average, the sample should look like the population
The standard approach for prevalence studies and surveys
Slide 3One way to achieve generalizability is through representative sampling. If the sample is drawn randomly from the target population, or drawn using probability methods with known selection probabilities, then the sample should, on average, look like the population. This is the standard approach for prevalence studies and surveys: sample representatively, weight appropriately, and the estimate generalizes.
Trials almost never work this way
Randomized trials randomly assign treatment; they rarely randomly sample participants
They enroll volunteers who meet eligibility criteria, consent, and can adhere to the protocol
That reflects the realities of conducting ethical research
Slide 4But trials almost never work this way. Randomized trials randomly assign treatment. They rarely randomly sample participants from the target population. Instead, trials enroll volunteers who meet specific eligibility criteria, provide informed consent, and can adhere to the study protocol. That reflects the realities of conducting ethical research. You cannot randomly select people and compel them to participate in a trial. Enrollment requires willingness, and willingness is itself selective.
Trial samples routinely differ from the populations they aim to inform
Participants tend to be younger, healthier, more educated, more engaged with the health system
Academic medical centers draw from different populations than community clinics
Strict eligibility excludes comorbidities, complex regimens, unstable circumstances
Slide 5As a result, trial samples routinely differ from the populations they aim to inform. Participants tend to be younger, healthier, more educated, and more engaged with the health system than typical patients. Trials conducted at academic medical centers draw from different populations than community clinics. And strict eligibility criteria exclude patients with comorbidities, complex medication regimens, or unstable social circumstances — exactly the patients who may later receive the treatment in routine care.
For trials, generalizability is an effect modification problem
Not the sampling question
Did we sample representatively?
We almost certainly did not
The question that decides it
Do the differences modify the effect?
If they don't, the estimate applies
Slide 6This means generalizability for trials is not primarily a sampling problem. It's an effect modification problem. The question is not whether we sampled representatively — we almost certainly didn't. The question is whether the characteristics on which our sample differs from the target population modify the treatment effect. If they don't, the estimate still applies. If they do, it may not.
2011 Nepal DHS — design-based
Representative sampling makes generalizability easier
Multistage cluster sampling across Nepal
10,826 households; 12,674 women ages 15 to 49
13 eco-development regions, each stratified by urban and rural areas
Sampling weights calculated to account for the complex design
Slide 7The 2011 Demographic and Health Survey illustrates the clearest path to generalizability. Researchers used multistage cluster sampling to survey more than ten thousand households nationwide, interviewing nearly thirteen thousand women ages fifteen to forty-nine. The sampling design was explicitly constructed to provide nationally representative estimates: thirteen eco-development regions, each stratified by urban and rural areas, with sampling weights calculated to account for the complex design.
2011 Nepal DHS — design-based
The estimates apply to Nepali women, not only to the women surveyed
What the survey found
50% of married women using a method
27% with unmet need for family planning
Why it generalizes
Sample drawn from the target population
Probability methods, known selection
Weights adjust for the design
Slide 8The target population was clearly defined — women of reproductive age in the country — and the study sample was designed to be a representative subset of it. The survey could estimate that fifty percent of currently married women were using some method of contraception, and that twenty-seven percent had an unmet need for family planning. Because the sample was drawn from the target population using probability methods, and because weights adjust for known selection probabilities, these estimates generalize. They apply to women of reproductive age across the country, not only to the women surveyed.
The gold standard for descriptive generalizability
Design the sample to represent the target
Slide 9That is the gold standard for descriptive generalizability. Design the sample to represent the target, use probability methods, and adjust for selection. The inference is carried by the design itself.
Hypothetical — convenience sample
Many studies have no probability sample to work with
Patients at a clinic, respondents to an online survey, people recruited through social networks
A team wants modern contraceptive use in Nepal before the next DHS is scheduled
An online survey promoted through social media collects 5,000 responses in weeks
Slide 10But what if you don't have a probability sample? Many studies rely on convenience samples — patients at a clinic, respondents to an online survey, or participants recruited through social networks. Consider a hypothetical example. A research team wants to estimate modern contraceptive use among women ages fifteen to forty-nine, before the next Demographic and Health Survey is scheduled. Instead of a costly household survey, they run an online survey promoted through social media and health-related websites. Within weeks, they collect responses from five thousand women across the country.
Hypothetical — convenience sample
The convenience is obvious; so is the problem
Respondents are overwhelmingly urban and younger
More educated, wealthier, more likely to have smartphone access
The raw estimate would almost certainly overstate national use
Slide 11The convenience is obvious, and so is the problem. Women who respond to an online survey are not representative of all women in the country. They are overwhelmingly urban, younger, more educated, wealthier, and more likely to have smartphone access. The raw estimate of modern contraceptive use from this sample would almost certainly overstate use in the national population.
MRP — model-based
Multilevel regression with poststratification, in two steps
Model the outcome on characteristics present in both the sample and a population reference
Age group, education, urban or rural residence, wealth quintile
Then reweight the model's predictions to the population's composition
Slide 12One approach to this is multilevel regression with poststratification. The logic works in two steps. First, fit a model that predicts the outcome — contraceptive use — from respondent characteristics available in both the online survey and a population reference, such as census data or a nationally representative survey. Age group, education level, urban or rural residence, wealth quintile. The model estimates how contraceptive use varies across these groups, even where some groups are thin in the sample. Second, use population data to determine how common each combination of characteristics is in the target population, and weight the model's predictions accordingly. The result is an estimate weighted to the composition of the national population.
Wang et al., 2015
The method has worked on badly skewed samples
Survey data from Xbox users — 65% young men
Used to predict US presidential election outcomes
Accuracy comparable to traditional polls after poststratification
Slide 13The method has worked on badly skewed samples. In a well-known example, researchers used survey data collected from Xbox users — sixty-five percent of them young men — to predict United States presidential election outcomes, and achieved accuracy comparable to traditional polls after poststratification. The adjustment doesn't eliminate all bias, but it can reduce it dramatically when the adjustment variables capture the key differences between sample and population.
The critical assumption is that the model includes the variables that matter
The adjustment variables have to capture the differences that matter
Unmeasured differences survive the reweighting
Exposure to family planning information; autonomy over reproductive decisions
Slide 14The critical assumption is that the model includes the variables that matter. If contraceptive use depends heavily on characteristics that differ between online respondents and the general population, and those characteristics aren't captured in the model, then adjustment won't fully correct the bias. Suppose women who respond to online surveys differ from other women in their exposure to family planning information, or in their autonomy over reproductive decisions. If those factors go unmeasured, the adjusted estimate is still biased.
Two routes to generalizability, carrying different assumptions
Design-based
Probability sampling, known selection
The design carries the inference
Remains the gold standard
Model-based
Convenience sample, explicit model
The assumptions carry the inference
For when sampling isn't feasible
Slide 15So generalizability from a convenience sample is possible, but it requires explicit modeling assumptions that probability sampling does not. Under the design-based route, the sampling scheme carries the inference. Under the model-based route, the assumptions do. Multilevel regression with poststratification is valuable when representative sampling isn't feasible, but they trade design-based inference for model-based inference, with everything that entails.
Bhatt et al., 2021 — qualitative
Sometimes generalizability is simply not achievable
Focus group discussions and in-depth interviews with people ages 15 to 24 in Nepal
All data collection in a single village — Hattimuda, in Morang district
Perceived barriers, gender dynamics, the gap between what schools teach and what is understood
Slide 16Sometimes generalizability is simply not achievable, and the honest response is to acknowledge the limits of inference. Consider a qualitative study examining perceptions of family planning services among young people. The researchers conducted focus group discussions and in-depth interviews with adolescents and young adults ages fifteen to twenty-four. All data collection occurred in a single village — Hattimuda, in Morang district. The study reveals how young people there think about family planning: the barriers they perceive, the role of gender dynamics, and the gap between what schools teach and what young people actually understand.
Bhatt et al., 2021 — qualitative
The authors state the limit directly
"Our findings might differ if the sample had been drawn from other parts of the country."
Young people in one eastern village may not share the perceptions of youth in Kathmandu
Slide 17That is valuable knowledge. But young people recruited in Hattimuda are not a representative subset of the country's youth in any formal sense. Young people in one eastern village may or may not share the perceptions of youth in urban Kathmandu, the Terai plains, or the western hill districts. The authors acknowledge this directly, writing that their findings might differ if the sample had been drawn from other parts of the country. The findings provide real insight into one context; extending them to the national level requires assumptions the study itself cannot verify.
Hypothetical — causal
The causal version of the same question
A randomized trial of a community health worker intervention in Nepal
20 primary health centers across four districts; 2,000 women not using modern methods
After 12 months, a 15 percentage point increase in uptake over the control arm
Slide 18Now consider a causal question. Suppose researchers conduct a randomized trial evaluating a community health worker intervention to increase modern contraceptive use among married women. The trial runs in twenty primary health centers across four districts, enrolling two thousand women who are not currently using modern methods. After twelve months, women in the intervention arm show a fifteen percentage point increase in modern method uptake compared with the control arm.
Hypothetical — causal
Strong internal validity, an open generalizability question
Randomization makes the effect estimate unbiased for the women enrolled
The Ministry of Health wants to know whether the effect would hold if scaled nationally
Trial participants are a subset of Nepali women of reproductive age
Slide 19The trial has strong internal validity. Randomization ensures that the effect estimate is unbiased for the women who enrolled. But the Ministry of Health wants to know whether this effect would hold if the intervention were scaled nationally. That is a generalizability question. The trial participants are a subset of all women of reproductive age in the country, but they may not represent that larger population.
Hypothetical — causal
Who the trial enrolled, and who it missed
Recruited at health centers, so already engaged with the health system
The four districts may differ in urbanization, ethnicity, or infrastructure
Women who never visit a health center, or face family opposition, are missing
Slide 20And indeed, they probably don't. Trial participants were recruited from health centers, meaning they were already engaged with the health system. They consented to a research study, suggesting openness to new information. The four districts may differ from others in urbanization, ethnicity, or health infrastructure. And the women with the most barriers to contraceptive use — those who never visit a health center, who face stronger family opposition, or who live in remote areas — are systematically underrepresented.
For causal effects the concern takes a specific form
If the effect is constant
Works equally well across groups
The trial estimate applies nationally
Even with an unrepresentative sample
If the effect varies
Works better where women are engaged
Easier to reach, more receptive
15 points may overstate what scales
Slide 21The generalizability concern is familiar from the descriptive examples, but for causal effects it takes a specific form. What matters is not whether the trial sample differs from the national population on any characteristic, but whether it differs on characteristics that modify the treatment effect. If the intervention works equally well regardless of health-seeking behavior, education, or geographic isolation, then the trial estimate applies nationally even though the sample was unrepresentative. But if the intervention works better among women already engaged with health services — because they're easier to reach, more receptive to counseling, or face fewer structural barriers — then the fifteen percentage point effect may overstate what would happen at scale.
In Closing
Only differences that modify the estimate threaten generalizability
Whether the sample was built by design, repaired by model, or acknowledged as local, the test is the same
Slide 22Four studies, and four different things each can claim: an estimate carried by its sampling design, an estimate carried by a model's assumptions, findings honest about belonging to one village, and a causal effect whose reach depends on whether the intervention works the same way everywhere. Whichever you have in front of you, the test is the same. Do the sample and the target differ on characteristics that modify the estimate?