Video 2 of 4

Does the sample stand for the population it came from?

Generalizability asks whether results from your sample stand for the population that sample came from, worked four times: an estimate carried by its sampling design, one carried by a model’s assumptions, findings honest about belonging to one village, and a causal effect whose reach depends on effect modification.

10:14 · 22 slides · printable slides · transcript

Slides

Printable deck →

Does the sample stand for the population it came from?

1 / 22
Transcript22 sections

Generated from the narration script. Plain text version.

1Generalizability is the first of the two external validity problems. It asks whether results from a study sample apply to the target population that sample was drawn from. Its defining characteristic is containment: everyone in the study sample is also a member of the target population.

2That containment is what separates this problem from the one in the next video. In generalizability, everyone we studied belongs to the population we want to inform. In transportability, they don't. Everything that follows here assumes the first arrangement.

3One way to achieve generalizability is through representative sampling. If the sample is drawn randomly from the target population, or drawn using probability methods with known selection probabilities, then the sample should, on average, look like the population. This is the standard approach for prevalence studies and surveys: sample representatively, weight appropriately, and the estimate generalizes.

4But trials almost never work this way. Randomized trials randomly assign treatment. They rarely randomly sample participants from the target population. Instead, trials enroll volunteers who meet specific eligibility criteria, provide informed consent, and can adhere to the study protocol. That reflects the realities of conducting ethical research. You cannot randomly select people and compel them to participate in a trial. Enrollment requires willingness, and willingness is itself selective.

5As a result, trial samples routinely differ from the populations they aim to inform. Participants tend to be younger, healthier, more educated, and more engaged with the health system than typical patients. Trials conducted at academic medical centers draw from different populations than community clinics. And strict eligibility criteria exclude patients with comorbidities, complex medication regimens, or unstable social circumstances — exactly the patients who may later receive the treatment in routine care.

6This means generalizability for trials is not primarily a sampling problem. It's an effect modification problem. The question is not whether we sampled representatively — we almost certainly didn't. The question is whether the characteristics on which our sample differs from the target population modify the treatment effect. If they don't, the estimate still applies. If they do, it may not.

7The 2011 Demographic and Health Survey illustrates the clearest path to generalizability. Researchers used multistage cluster sampling to survey more than ten thousand households nationwide, interviewing nearly thirteen thousand women ages fifteen to forty-nine. The sampling design was explicitly constructed to provide nationally representative estimates: thirteen eco-development regions, each stratified by urban and rural areas, with sampling weights calculated to account for the complex design.

8The target population was clearly defined — women of reproductive age in the country — and the study sample was designed to be a representative subset of it. The survey could estimate that fifty percent of currently married women were using some method of contraception, and that twenty-seven percent had an unmet need for family planning. Because the sample was drawn from the target population using probability methods, and because weights adjust for known selection probabilities, these estimates generalize. They apply to women of reproductive age across the country, not only to the women surveyed.

9That is the gold standard for descriptive generalizability. Design the sample to represent the target, use probability methods, and adjust for selection. The inference is carried by the design itself.

10But what if you don't have a probability sample? Many studies rely on convenience samples — patients at a clinic, respondents to an online survey, or participants recruited through social networks. Consider a hypothetical example. A research team wants to estimate modern contraceptive use among women ages fifteen to forty-nine, before the next Demographic and Health Survey is scheduled. Instead of a costly household survey, they run an online survey promoted through social media and health-related websites. Within weeks, they collect responses from five thousand women across the country.

11The convenience is obvious, and so is the problem. Women who respond to an online survey are not representative of all women in the country. They are overwhelmingly urban, younger, more educated, wealthier, and more likely to have smartphone access. The raw estimate of modern contraceptive use from this sample would almost certainly overstate use in the national population.

12One approach to this is multilevel regression with poststratification. The logic works in two steps. First, fit a model that predicts the outcome — contraceptive use — from respondent characteristics available in both the online survey and a population reference, such as census data or a nationally representative survey. Age group, education level, urban or rural residence, wealth quintile. The model estimates how contraceptive use varies across these groups, even where some groups are thin in the sample. Second, use population data to determine how common each combination of characteristics is in the target population, and weight the model's predictions accordingly. The result is an estimate weighted to the composition of the national population.

13The method has worked on badly skewed samples. In a well-known example, researchers used survey data collected from Xbox users — sixty-five percent of them young men — to predict United States presidential election outcomes, and achieved accuracy comparable to traditional polls after poststratification. The adjustment doesn't eliminate all bias, but it can reduce it dramatically when the adjustment variables capture the key differences between sample and population.

14The critical assumption is that the model includes the variables that matter. If contraceptive use depends heavily on characteristics that differ between online respondents and the general population, and those characteristics aren't captured in the model, then adjustment won't fully correct the bias. Suppose women who respond to online surveys differ from other women in their exposure to family planning information, or in their autonomy over reproductive decisions. If those factors go unmeasured, the adjusted estimate is still biased.

15So generalizability from a convenience sample is possible, but it requires explicit modeling assumptions that probability sampling does not. Under the design-based route, the sampling scheme carries the inference. Under the model-based route, the assumptions do. Multilevel regression with poststratification is valuable when representative sampling isn't feasible, but they trade design-based inference for model-based inference, with everything that entails.

16Sometimes generalizability is simply not achievable, and the honest response is to acknowledge the limits of inference. Consider a qualitative study examining perceptions of family planning services among young people. The researchers conducted focus group discussions and in-depth interviews with adolescents and young adults ages fifteen to twenty-four. All data collection occurred in a single village — Hattimuda, in Morang district. The study reveals how young people there think about family planning: the barriers they perceive, the role of gender dynamics, and the gap between what schools teach and what young people actually understand.

17That is valuable knowledge. But young people recruited in Hattimuda are not a representative subset of the country's youth in any formal sense. Young people in one eastern village may or may not share the perceptions of youth in urban Kathmandu, the Terai plains, or the western hill districts. The authors acknowledge this directly, writing that their findings might differ if the sample had been drawn from other parts of the country. The findings provide real insight into one context; extending them to the national level requires assumptions the study itself cannot verify.

18Now consider a causal question. Suppose researchers conduct a randomized trial evaluating a community health worker intervention to increase modern contraceptive use among married women. The trial runs in twenty primary health centers across four districts, enrolling two thousand women who are not currently using modern methods. After twelve months, women in the intervention arm show a fifteen percentage point increase in modern method uptake compared with the control arm.

19The trial has strong internal validity. Randomization ensures that the effect estimate is unbiased for the women who enrolled. But the Ministry of Health wants to know whether this effect would hold if the intervention were scaled nationally. That is a generalizability question. The trial participants are a subset of all women of reproductive age in the country, but they may not represent that larger population.

20And indeed, they probably don't. Trial participants were recruited from health centers, meaning they were already engaged with the health system. They consented to a research study, suggesting openness to new information. The four districts may differ from others in urbanization, ethnicity, or health infrastructure. And the women with the most barriers to contraceptive use — those who never visit a health center, who face stronger family opposition, or who live in remote areas — are systematically underrepresented.

21The generalizability concern is familiar from the descriptive examples, but for causal effects it takes a specific form. What matters is not whether the trial sample differs from the national population on any characteristic, but whether it differs on characteristics that modify the treatment effect. If the intervention works equally well regardless of health-seeking behavior, education, or geographic isolation, then the trial estimate applies nationally even though the sample was unrepresentative. But if the intervention works better among women already engaged with health services — because they're easier to reach, more receptive to counseling, or face fewer structural barriers — then the fifteen percentage point effect may overstate what would happen at scale.

22Four studies, and four different things each can claim: an estimate carried by its sampling design, an estimate carried by a model's assumptions, findings honest about belonging to one village, and a causal effect whose reach depends on whether the intervention works the same way everywhere. Whichever you have in front of you, the test is the same. Do the sample and the target differ on characteristics that modify the estimate?

GeneralizabilityProbability SamplingPoststratificationSurvey Design