From efficacy to impact: translation, implementation, and evaluation Chapter 1: Global Health Research — Video 4 https://ghrbook.com/videos/from-efficacy-to-impact/ [Slide 1] This video follows the road from a good idea to real impact. We look at where ideas stall on the way to scale, at the fields that study those bottlenecks, and at monitoring and evaluation. [Slide 2] It's a long road from idea to impact, and most ideas don't complete the trip. Why does so much promising research fail to reach the people who need it? Practitioners of translational research point to four key bottlenecks, four places along the way where ideas often stall. They're labeled T1 through T4. [Slide 3] T1 is translation from basic science to clinical research, the move from bench to bedside. T2 is translation from early clinical trials to Phase 3 trials and beyond, with larger patient populations. Later we recognized two more. T3 is translation from efficacy trials, meaning Phase 3 trials, to real-world effectiveness through implementation research. And T4 is translation from evidence about delivery at scale to the adoption of new policies. It's these last two that keep good ideas from reaching population health and policy. [Slide 4] Here's the map from the last video, with the translation arrows in view. T1 runs from basic research into the clinical pipeline. T2 and T3 sit inside that pipeline, moving from early trials through Phase 3 and out toward effectiveness. And T4 runs off the right-hand side, from evidence about delivery to the adoption of policy, and to scaling up what works to improve population health. [Slide 5] This is an example of what we call a delivery gap, or a know-do gap, in global health. It's the space between discovering what works and delivering the solution at scale. This figure plots years from launch along the horizontal axis and global coverage on the vertical. The three dashed curves are United States averages, for vaccines, for drugs, and for diagnostics. They climb to full coverage within about five years. Every solid line is a product being delivered globally, and the picture there is completely different. Oral rehydration therapy, or ORT, highlighted in the legend, launched in 1971, and nearly thirty years on its curve is still climbing. [Slide 6] Policy and implementation research aims to close these gaps. It has been called the science of scale up. The gap between what we know and what we do is itself a research problem, and this is the field that takes it on. It combines two things: implementation research, and health policy and systems research. Let's take them one at a time. [Slide 7] Implementation research is the study of strategies for expanding the reach and coverage of tested ideas, so that they improve population health. The common scenario is the one we just saw. Intervention research has already demonstrated that a treatment is efficacious, ORT for example, but take-up is low and slow, which limits its impact. So the implementation research question is how best to promote its use. [Slide 8] Health policy and systems research shares that same goal of population impact, but it takes an even broader view, looking at the systemic and policy factors that can hinder or facilitate scale up. [Slide 9] Here's an example of what that looks like. A retrospective analysis evaluated the impact of policymaking in Uganda on ORT and zinc coverage. The authors triangulated data from various sources on government actions, the distribution of supplies, and treatment with ORT. They estimated that the proportion of young children with diarrhea who received ORT increased thirty-fold between 2011 and 2016, from 1% to 30%. And their policy analysis concluded that government actions likely made the difference. [Slide 10] Another area of applied work in global health is monitoring and evaluation, also known as program evaluation. Whether it counts as research is a question with a long history, so let's start there. [Slide 11] Program evaluation became commonplace in the United States by the end of the nineteen fifties, and it grew dramatically in the nineteen sixties as the federal government expanded and introduced new social programs. Lawmakers wanted accountability, and the evaluation of social programs took off. [Slide 12] But is program evaluation really research? Methods giant Donald Campbell thought so. Writing in 1969, he argued that the United States should be ready for an experimental approach to social reform, an approach in which we try out new programs designed to cure specific problems, in which we learn whether or not these programs are effective, and in which we retain, imitate, modify or discard them on the basis of their effectiveness on the multiple imperfect criteria available. [Slide 13] But not everyone agrees. Some have argued that program evaluation is really designed for program implementers and funders, and that the messy nature of program implementation requires a loosening of research standards. [Slide 14] One standard introductory text on evaluation strikes a balance on this question. Their answer is perhaps a bit unsatisfying, but it is arguably true nevertheless: it depends. In essence, program evaluations should be as rigorous as logistics, ethics, politics, and resources permit, and no less. Some evaluations are more rigorous than others, and those will meet our definition of research: a systematic investigation designed to develop or contribute to generalizable knowledge. [Slide 15] That is why, on the map, I represent monitoring and evaluation as a block that intersects applied research but extends outside the research boundary. Some of what gets done under that heading meets the definition of research, and some of it sits outside, so the block straddles the line. [Slide 16] Evaluations can take different forms and serve various purposes. A subset of both applied research and program evaluation is the impact evaluation. An impact evaluation is a study that aims to quantify the causal effect, or impact, of a program or policy on some outcome of interest. There are several research designs that can generate evidence of impact. One example is the randomized controlled trial, a mainstay of Phase 2 and Phase 3 clinical trials. Not all impact evaluations use random assignment to make causal inferences, but they all share the goal of making a cause-and-effect claim. And as we'll see in later chapters, the strength of that claim rests on the assumptions of the particular research design. [Slide 17] Evaluation and monitoring answer two different questions. Evaluation asks: do they work? Monitoring asks: what happened? Program monitoring is concerned with documenting the implementation of programs and interventions. How are resources being used? How many people participate? Does the program reach the intended targets? Not all programs are evaluated, but most are monitored to some degree, for accountability to funders. Researchers can also use monitoring data to document participants' exposure to the program, and to conduct economic analyses related to program costs and impacts. [Slide 18] A related activity is the process evaluation. A process evaluation goes beyond monitoring counts and tallies, largely an administrative task, to ask if a program is being delivered as intended. This gets at the question of fidelity of the implementation to the original design. [Slide 19] Process evaluations are essential for impact evaluations. If a program fails to show an impact, the next question is: why? Did the program fail because the idea or theory behind the program was wrong? That's theory failure. Or was the implementation of the program so troubled that there was never a chance for success? That's implementation failure.