Difference-in-differences and maternal health vouchers in Bangladesh Chapter 11: Quasi-Experimental Designs — Video 3 https://ghrbook.com/videos/difference-in-differences/ [Slide 1] Difference-in-differences combines the best features of pre-post and between-group designs while mitigating some weaknesses of each. It has become a workhorse method in policy evaluation and program impact assessment. [Slide 2] We measure outcomes in both treatment and comparison groups, both before and after the intervention. We calculate how much the treatment group changed, and how much the comparison group changed over the same period. The difference between those two changes is our estimate of the treatment effect. [Slide 3] The comparison group is doing the heavy lifting. Its change over time estimates what would have happened to the treatment group without the intervention: the background trend from seasonal variation, economic shifts, or anything else that affected both groups equally. Subtracting it out leaves us with the treatment effect, in principle. Unlike pre-post, the design accounts for time trends, and unlike post-only comparisons, it accounts for baseline differences between groups. [Slide 4] Starting in 2006, the government of Bangladesh rolled out the Maternal Health Voucher Scheme across dozens of subdistricts, the basic unit of local governance. The program gave pregnant women vouchers they could exchange for antenatal care, skilled delivery, and postnatal services at public or private providers. The goal was to reduce financial barriers, increase facility-based deliveries, and ultimately reduce maternal and neonatal deaths. [Slide 5] The program wasn't randomized. Subdistricts were selected based on poverty, literacy, and the presence of health workers to administer the program, so by design, treated areas were poorer and more rural than untreated ones. Before the voucher scheme, only 7.2% of women in treated subdistricts delivered in a health facility, compared to 15.3% in control areas. That's the kind of baseline imbalance that makes a simple post-treatment comparison unreliable. [Slide 6] The evaluation drew on five rounds of Bangladesh Demographic and Health Survey data, covering births from 2000 to 2016. The treatment group was the 55 subdistricts that received the voucher scheme, and the comparison group was roughly 456 subdistricts that didn't. Because the program rolled out in stages over about four years, the authors used event study models that estimated separate effects for each two-year period before and after implementation, so they could assess both parallel trends and the timing of any effects. [Slide 7] The parallel trends evidence was strong for most outcomes. These are the raw trends in treated and control subdistricts from 2000 to 2016. For institutional delivery, the two lines track closely through the pre-intervention years, then begin to diverge after the voucher scheme starts, which is what a credible difference-in-differences analysis should look like. [Slide 8] The event study coefficients tell the same story more formally. Pre-intervention estimates hover near zero, confirming that treated and control subdistricts were on similar trajectories before the program. Effects then emerge gradually, with a lag of two to four years, consistent with a program that takes time to change behavior. [Slide 9] After six years of access to the voucher scheme, the probability of delivering in a health facility increased by 6.5 percentage points, with a 95% confidence interval from negative 0.6 to 13.6. Having a skilled birth attendant increased by 5.8 percentage points, with an interval from negative 1.8 to 13.3. These are meaningful increases in a context where baseline facility delivery was below 10%. [Slide 10] Despite these gains in service utilization, the study found no evidence that the voucher program reduced stillbirth, neonatal mortality, or infant mortality. The estimates were small and imprecise, hovering near zero. More women were delivering in facilities, but babies weren't surviving at higher rates. [Slide 11] The authors point to several possibilities. The program may not have reached the highest-risk women: eligibility criteria weren't always enforced, and awareness was uneven. Providers weren't adequately compensated, potentially reducing quality of care as patient volume increased. And facility birth in Bangladesh is associated with lower breastfeeding initiation, which could offset some of the survival benefits of skilled attendance. [Slide 12] The design can tell us that the program changed behavior, but it can't tell us why that change didn't translate into the outcome that mattered most. It also can't resolve whether the null mortality finding reflects true program failure, insufficient statistical power, or confounding from concurrent changes in treated areas. The authors are transparent about each of these possibilities, which is what good quasi-experimental research looks like.