Chapter 11 · Video 2

Comparing groups without baseline data

11 slides · Video page · All videos · Transcript
Print layout
Slide 1

Comparing groups without baseline data

Slide 2

Comparing two groups at the same point in time

One group got the intervention and one didn't
If the groups are comparable, the difference is the treatment effect
Here, nobody randomized anything
Slide 3
Multiple group post-test only
The comparison removes history effects
Multiple group post-test only design. Reproduced from Chapter 11.
Slide 4

Two villages, two malaria prevention strategies

Village A
Insecticide-treated bed nets
Plus indoor residual spraying
40% lower incidence three months later
Village B
Insecticide-treated bed nets only
Slide 5

Why might Village A have lower incidence anyway?

It was selected because it had better health infrastructure
Its residents were more health-conscious
It's at higher elevation, with fewer mosquitoes
Slide 6
Key assumption

The groups were similar enough before treatment

The comparison group's outcome stands in for the counterfactual
Without random assignment, we don't know the groups were comparable
Slide 7
Mitchell et al., 2018

The Millennium Villages Project

Launched in 2005 across ten sub-Saharan African countries
Agriculture, nutrition, education, health, water and sanitation
No randomized evaluation, no comparison villages at the outset
Slide 8
Mitchell et al., 2018

Choosing comparison villages ten years after the project began

Matched retrospectively on survey data, GIS, and wealth, education and health indices
Project and comparison villages surveyed in 2015
Outcomes compared across 40 indicators
Slide 9
Mitchell et al., 2018

What the endline evaluation found

30 of 40 impact estimates significant, all favoring the project
Strongest effects in agriculture and health
Under-5 mortality 23 per 1,000 livebirths lower (95% UI 6 to 40)
Poverty, nutrition and education less conclusive
Slide 10

What the comparison cannot show

Sites selected for undernutrition, agroecological diversity and political buy-in
No baseline data from the comparison villages
Earlier reports presented before-after differences as “impacts”
Slide 11
In Closing

Matching can't prove the groups were equivalent before treatment

A comparison group and rigorous matching improved on earlier evaluations
The fundamental limitation remains

Comparing groups without baseline data

Slide 1The pre-post design's fatal flaw is that we can't easily separate the intervention from everything else that changed over the same period. One way around this is to compare two different groups at the same point in time.

Comparing two groups at the same point in time

One group got the intervention and one didn't
If the groups are comparable, the difference is the treatment effect
Here, nobody randomized anything
Slide 2One group got the intervention, and one didn't. If the groups are comparable, any difference in outcomes is our estimate of the treatment effect. That's a big if. In a randomized trial, randomization makes the groups comparable by design. Here, nobody randomized anything.
Multiple group post-test only
The comparison removes history effects
Multiple group post-test only design. Reproduced from Chapter 11.
Slide 3By comparing groups at the same point in time, we eliminate the history effects that plague pre-post designs. Both groups experienced the same historical events, seasonal patterns, and broader social changes. But we've introduced a new problem: selection bias. How did the groups end up in treatment or control? If assignment wasn't random, the groups might differ in ways that affect the outcome, and those differences get confused with the treatment effect.

Two villages, two malaria prevention strategies

Village A
Insecticide-treated bed nets
Plus indoor residual spraying
40% lower incidence three months later
Village B
Insecticide-treated bed nets only
Slide 4Consider evaluating two malaria prevention strategies in different villages. Village A received insecticide-treated bed nets plus indoor residual spraying. Village B received only bed nets. We measure malaria incidence three months later and find Village A has 40% lower incidence.

Why might Village A have lower incidence anyway?

It was selected because it had better health infrastructure
Its residents were more health-conscious
It's at higher elevation, with fewer mosquitoes
Slide 5Was it the combination intervention? Maybe. But what if Village A was selected for the intensive intervention because it had better health infrastructure, making implementation easier? What if Village A residents were more health-conscious, which is why they were chosen, and would have had lower malaria incidence anyway? Or what if Village A is at higher elevation, with fewer mosquitoes?
Key assumption

The groups were similar enough before treatment

The comparison group's outcome stands in for the counterfactual
Without random assignment, we don't know the groups were comparable
Slide 6The comparison group's outcome stands in for what the treatment group would have experienced without the intervention. The key assumption is that the two groups were similar enough before treatment that any post-treatment difference reflects the intervention, and not pre-existing differences. Without random assignment, we don't know if the groups were comparable before the intervention.
Mitchell et al., 2018

The Millennium Villages Project

Launched in 2005 across ten sub-Saharan African countries
Agriculture, nutrition, education, health, water and sanitation
No randomized evaluation, no comparison villages at the outset
Slide 7The Millennium Villages Project is one of the most ambitious, and most debated, development interventions of this century. Launched in 2005 across ten sub-Saharan African countries, the project poured integrated investments into rural villages: agriculture, nutrition, education, health infrastructure, water and sanitation. The goal was to achieve the Millennium Development Goals within a decade. Hundreds of millions of dollars were spent, but no randomized evaluation was planned, and no comparison villages were selected at the outset.
Mitchell et al., 2018

Choosing comparison villages ten years after the project began

Matched retrospectively on survey data, GIS, and wealth, education and health indices
Project and comparison villages surveyed in 2015
Outcomes compared across 40 indicators
Slide 8By the time the endline evaluation was designed, the project had been running for ten years. That's the most common reason for a post-test only design: the intervention started before the evaluation was designed. The researchers couldn't go back in time to collect baseline data from comparison villages, so they retrospectively selected matched villages using Demographic and Health Survey data, geographic information systems, and indices of wealth, education, and health. They surveyed both project and comparison villages in 2015 and compared outcomes across 40 indicators.
Mitchell et al., 2018

What the endline evaluation found

30 of 40 impact estimates significant, all favoring the project
Strongest effects in agriculture and health
Under-5 mortality 23 per 1,000 livebirths lower (95% UI 6 to 40)
Poverty, nutrition and education less conclusive
Slide 9The results were largely favorable. Impact estimates for 30 of 40 outcomes were significant, all favoring the project villages. The strongest effects were in agriculture and health, consistent with where the project invested most. Mortality among children under five was 23 deaths per 1,000 livebirths lower in project villages, with a 95% uncertainty interval of 6 to 40. But effects on poverty, nutrition, and education were less conclusive.

What the comparison cannot show

Sites selected for undernutrition, agroecological diversity and political buy-in
No baseline data from the comparison villages
Earlier reports presented before-after differences as “impacts”
Slide 10The study is honest about its limitations. Project sites were selected for high undernutrition, agroecological diversity, and local political buy-in, all factors that could confound the comparison. The matched villages looked similar on measured characteristics, but without baseline data from before the project started, there's no way to verify they were truly comparable ten years earlier. Previous evaluations had been criticized for exactly this problem: earlier reports presented before-after differences within project villages as impacts, conflating pre-post change with causal effects.
In Closing

Matching can't prove the groups were equivalent before treatment

A comparison group and rigorous matching improved on earlier evaluations
The fundamental limitation remains
Slide 11This endline evaluation improved on those efforts by adding a comparison group and using rigorous matching, but the fundamental limitation remains. A post-test comparison, no matter how carefully matched, can't prove that groups were equivalent before treatment.