Video 6 of 6
Instrumental variables, and choosing a design
An instrument changes who gets treated without affecting the outcome any other way, and that exclusion restriction is the part that usually fails. Encouragement designs, compliers and never-takers, monotonicity and the local average treatment effect, a randomized campaign that helped households in Tangiers get piped water, and a closing guide to choosing among the six designs.
8:51 · 19 slides · printable slides · transcript
Slides
Printable deck →▶Transcript19 sections
Generated from the narration script. Plain text version.
1Instrumental variables are the last of the six designs in this chapter. After the Tangiers water study, we'll put all six side by side.
2Consider trying to estimate the effect of education on health. People who get more education are different from those who don't. They might have higher cognitive ability, more supportive families, better health to begin with, or greater motivation, and these factors affect both education and health. Even with rich data, we probably can't measure all these confounders.
3Instrumental variables offer a solution: find something that affects treatment but affects the outcome only through its effect on treatment. In the education example, researchers have used changes in compulsory schooling laws as an instrument. When a country raises the minimum school-leaving age from 14 to 16, some people get more education than they otherwise would have because the law changed, and not because of their motivation or family background. If the law affects health only by changing how long people stay in school, the variation it creates is the clean variation this design needs.
4The counterfactual here compares what happens when the instrument pushes more people toward treatment with what happens when it pushes fewer. The outcome difference between instrument groups, scaled by how much the instrument shifts treatment uptake, reveals the causal effect, but only for the people the instrument actually moved, the compliers. The key assumption is that the instrument affects the outcome only by changing treatment. If the instrument has a back door to the outcome, the whole strategy collapses.
5In a standard randomized trial, we randomize treatment assignment and hope everyone complies. In practice, compliance is rarely perfect. Some people assigned to treatment don't take it up, and some people assigned to control find a way to get treated anyway. An encouragement design embraces this reality. We randomize encouragement, like an invitation, a subsidy, or a home visit, and use the randomized encouragement as an instrument for actual treatment.
6This is instrumental variables with a randomized instrument. Randomization guarantees exchangeability and makes the exclusion restriction more defensible, though not automatic. The encouragement must affect the outcome only through its effect on treatment uptake. A home visit that encourages vaccination but also provides health education affects health through two channels, and that violates exclusion.
7In many trials, noncompliance runs in both directions, which creates four types of people. Compliers take treatment when encouraged and don't when not. Always-takers take treatment regardless, never-takers refuse regardless, and defiers do the opposite of their assignment. The always-takers and never-takers provide no information about the effect of treatment, because the instrument didn't change their behavior. That's why the design estimates a local average treatment effect: the effect among the people the instrument actually moved.
8For that estimate to be interpretable, we need a monotonicity assumption: the instrument must push everyone in the same direction. If encouragement makes some people more likely to take treatment and others less likely, the estimate becomes an uninterpretable weighted average. In most settings, monotonicity is reasonable. It's hard to imagine why being offered help with paperwork would cause some households to avoid connecting to the water system. But where the instrument could plausibly cut both ways, the assumption deserves scrutiny.
9Whether the local average treatment effect is useful depends on the policy question. If we want to know whether to expand a voluntary program, and the instrument is program availability, the effect among compliers is exactly what we need, because they're the people the expansion would reach. If we want to know what would happen if everyone received treatment, it's less helpful, because it says nothing about always-takers or never-takers.
10In the low-income neighborhoods of Tangiers, Morocco, roughly 845 households lacked a private water connection and couldn't afford the fee to get piped water into their homes. Most relied on public taps or bought water from neighbors. The utility company offered interest-free credit to cover the cost, but take-up was low: only about 10% of eligible households applied on their own. The barriers weren't just financial. The application process required obtaining authorization from local authorities, providing photocopies of identification documents, and making a down payment at a branch office.
11A research team designed a randomized encouragement trial. They randomly assigned 434 households to a door-to-door campaign that explained the credit program and handled the paperwork on the spot, while 411 households served as controls, eligible for the same credit but without the individualized assistance. The randomized encouragement is the instrument, and actual connection to piped water is the treatment. By six months, 69% of encouraged households had purchased a connection, compared to 10% of controls.
12The results were surprising. Piped water connections did not improve child health. There was no reduction in waterborne disease, which makes sense: these households already had access to clean water through public taps. The water quality didn't change, only the convenience of access.
13But the connections improved well-being through a different channel. Among households that had been carrying containers to the public tap, about 43% of the sample, the time spent fetching water averaged over seven hours per week, time that a private connection eliminated. That time went to leisure and socializing, and not to labor market participation or schooling. Conflict with neighbors over water access dropped, and life satisfaction increased substantially.
14The study illustrates several features of good practice. The instrument is strong, with a first-stage difference of 59 percentage points. Exclusion is defensible, because the information campaign provided no health services, no income, and no goods other than help with paperwork. The researchers checked that treatment and control groups were balanced at baseline across 57 characteristics. And the findings challenge the assumption that water infrastructure improvements work primarily through health, an assumption that would have led to the wrong conclusion if the study had measured only waterborne disease.
15The trade-off is that we estimate the effect for compliers only. In this case, the compliers are households who would connect to the water system if someone helped them navigate the bureaucracy but wouldn't do it on their own. That's a policy-relevant group: they're exactly the people a simplified application process would reach.
16Here are the six designs side by side. Each makes different assumptions about what serves as the counterfactual, and each is vulnerable to different threats. A pre-post design assumes nothing else changed, and its main threats are history and maturation; it's best for exploratory work and formative evaluation. A post-test only comparison assumes the groups were comparable, and selection bias is the main threat; it's best for retrospective evaluation with no baseline data. Difference-in-differences assumes parallel trends, and differential trends and spillovers are the threats.
17An interrupted time series assumes the trend would have continued, and concurrent events are the threat; it's the design for population-level policies with no comparison group. Regression discontinuity assumes no sorting at the cutoff, and manipulation and bundled treatments are the threats. Instrumental variables rest on the exclusion restriction, and weak or invalid instruments are the threat.
18Before reaching for any of these, we should ask seriously whether randomization is possible. Sometimes what seems infeasible becomes feasible with creativity: stepped-wedge designs, encouragement randomization, or phased rollouts with waitlist controls. Randomization eliminates the assumptions this entire chapter is about. When it's truly off the table, the choice among quasi-experimental designs depends on what data we have, what structure the setting provides, and which assumptions we can most credibly defend.
19John Snow didn't have a name for what he did. He found an accident of infrastructure that created comparison groups nearly as good as randomization, and he used it. Today we have names for these designs, formal assumptions, and statistical machinery to go with them, but the core logic hasn't changed: find variation in treatment that is plausibly unrelated to the outcome, and use it to estimate a causal effect. Parallel trends, exclusion restrictions, and no manipulation at the cutoff are claims about the world that we have to argue for. So when we read or conduct this kind of research, we keep asking the question that ties every design together: what would have happened without the intervention? Our answer is the counterfactual, and our design is the argument for why that answer is credible.