Video 3 of 6
Difference-in-differences and maternal health vouchers in Bangladesh
Comparing the change over time in a treated group with the change in a comparison group, so the comparison group’s trend stands in for what would have happened anyway. Parallel trends and event-study estimates, worked through five rounds of DHS data on Bangladesh’s Maternal Health Voucher Scheme, which raised facility delivery with no evidence of lower stillbirth, neonatal or infant mortality.
4:43 · 12 slides · printable slides · transcript
Slides
Printable deck →▶Transcript12 sections
Generated from the narration script. Plain text version.
1Difference-in-differences combines the best features of pre-post and between-group designs while mitigating some weaknesses of each. It has become a workhorse method in policy evaluation and program impact assessment.
2We measure outcomes in both treatment and comparison groups, both before and after the intervention. We calculate how much the treatment group changed, and how much the comparison group changed over the same period. The difference between those two changes is our estimate of the treatment effect.
3The comparison group is doing the heavy lifting. Its change over time estimates what would have happened to the treatment group without the intervention: the background trend from seasonal variation, economic shifts, or anything else that affected both groups equally. Subtracting it out leaves us with the treatment effect, in principle. Unlike pre-post, the design accounts for time trends, and unlike post-only comparisons, it accounts for baseline differences between groups.
4Starting in 2006, the government of Bangladesh rolled out the Maternal Health Voucher Scheme across dozens of subdistricts, the basic unit of local governance. The program gave pregnant women vouchers they could exchange for antenatal care, skilled delivery, and postnatal services at public or private providers. The goal was to reduce financial barriers, increase facility-based deliveries, and ultimately reduce maternal and neonatal deaths.
5The program wasn't randomized. Subdistricts were selected based on poverty, literacy, and the presence of health workers to administer the program, so by design, treated areas were poorer and more rural than untreated ones. Before the voucher scheme, only 7.2% of women in treated subdistricts delivered in a health facility, compared to 15.3% in control areas. That's the kind of baseline imbalance that makes a simple post-treatment comparison unreliable.
6The evaluation drew on five rounds of Bangladesh Demographic and Health Survey data, covering births from 2000 to 2016. The treatment group was the 55 subdistricts that received the voucher scheme, and the comparison group was roughly 456 subdistricts that didn't. Because the program rolled out in stages over about four years, the authors used event study models that estimated separate effects for each two-year period before and after implementation, so they could assess both parallel trends and the timing of any effects.
7The parallel trends evidence was strong for most outcomes. These are the raw trends in treated and control subdistricts from 2000 to 2016. For institutional delivery, the two lines track closely through the pre-intervention years, then begin to diverge after the voucher scheme starts, which is what a credible difference-in-differences analysis should look like.
8The event study coefficients tell the same story more formally. Pre-intervention estimates hover near zero, confirming that treated and control subdistricts were on similar trajectories before the program. Effects then emerge gradually, with a lag of two to four years, consistent with a program that takes time to change behavior.
9After six years of access to the voucher scheme, the probability of delivering in a health facility increased by 6.5 percentage points, with a 95% confidence interval from negative 0.6 to 13.6. Having a skilled birth attendant increased by 5.8 percentage points, with an interval from negative 1.8 to 13.3. These are meaningful increases in a context where baseline facility delivery was below 10%.
10Despite these gains in service utilization, the study found no evidence that the voucher program reduced stillbirth, neonatal mortality, or infant mortality. The estimates were small and imprecise, hovering near zero. More women were delivering in facilities, but babies weren't surviving at higher rates.
11The authors point to several possibilities. The program may not have reached the highest-risk women: eligibility criteria weren't always enforced, and awareness was uneven. Providers weren't adequately compensated, potentially reducing quality of care as patient volume increased. And facility birth in Bangladesh is associated with lower breastfeeding initiation, which could offset some of the survival benefits of skilled attendance.
12The design can tell us that the program changed behavior, but it can't tell us why that change didn't translate into the outcome that mattered most. It also can't resolve whether the null mortality finding reflects true program failure, insufficient statistical power, or confounding from concurrent changes in treated areas. The authors are transparent about each of these possibilities, which is what good quasi-experimental research looks like.