From a trial result to policy Chapter 10: Randomized Controlled Trials — Video 7 https://ghrbook.com/videos/from-a-trial-result-to-policy/ [Slide 1] Two trial designs exist because of how programs actually reach people, and once a trial has a result, the work of turning it into policy and practice begins. [Slide 2] Standard RCTs keep distinct intervention and control groups throughout the study. But what if it's ethically or politically problematic to withhold the intervention from control clusters indefinitely? Or what if we simply can't turn the intervention on everywhere at once, because we have limited trainers, equipment, or funding, and a phased rollout is the only feasible option? [Slide 3] If we're willing to randomize the order in which clusters receive the intervention, a stepped-wedge design becomes possible. All clusters start in the control condition and then cross over to the intervention at randomly determined time points. By the end of the study everyone has received the intervention, but we've still been able to compare intervention and control conditions. [Slide 4] The Peru tuberculosis study from our video on designing a trial used this design. All 34 health care centers in the district eventually implemented active case finding for TB contacts, but they started at different, randomly determined times over 20 months. That let the researchers compare TB diagnosis rates in centers that had already begun active case finding with centers that hadn't yet started, while respecting the health department's need to eventually implement the program everywhere. [Slide 5] In a conventional parallel design, clusters are randomized once and stay in their assigned condition. The stepped-wedge design crosses clusters from control to intervention in sequence, and a variation adds transition periods, time for the intervention to be implemented before data collection resumes. The design works well when a program will be scaled up regardless, so the question is whether it works, when sequential rollout is logistically necessary, and when we can plausibly assume that underlying trends are similar across clusters. [Slide 6] Most trials we'll encounter are superiority trials. They test whether an intervention is better than a comparator. But sometimes the goal is to show that a new treatment is not meaningfully worse, perhaps because it's cheaper, easier to administer, has fewer side effects, or can be manufactured locally. That's called non-inferiority. [Slide 7] Global health is full of non-inferiority questions, like testing shorter TB treatment regimens against standard six-month therapy, or comparing community health worker delivery of a service against facility-based delivery. In each case, the question is whether the new approach is good enough that its practical advantages justify adoption. [Slide 8] A non-inferiority trial tests whether a new treatment's effect falls within a pre-specified margin of the standard treatment's effect. If the standard treatment reduces mortality by 30%, and we define a non-inferiority margin of 10 percentage points, we're asking: does the new treatment reduce mortality by at least 20%? [Slide 9] This figure shows the four possible outcomes. Everything hinges on where the confidence interval falls relative to two lines: the standard treatment's effect of 30%, and the non-inferiority margin, a reduction of at least 20%. Only when the entire confidence interval stays to the right of the margin can we claim non-inferiority. [Slide 10] Non-inferiority is a one-sided question: is the new treatment not too much worse? A related but distinct goal is equivalence. Equivalence trials ask a two-sided question: are two treatments close enough in effect that we can consider them interchangeable? They're most common in pharmaceutical regulation, where a generic drug manufacturer needs to show that its product performs the same as the branded version. [Slide 11] Now suppose the trial found that the intervention significantly improved the outcome we care about. Translating that finding into policy and practice is one of the hardest parts. Most RCTs report average treatment effects, the average difference in outcomes between intervention and control groups across everyone enrolled. But average effects can mask important heterogeneity in who benefits. [Slide 12] The tuberculous meningitis trial in Vietnam, which we met in our video on control conditions, illustrates this. The primary analysis across all 120 participants didn't show a statistically significant benefit from aspirin. But planned subgroup analyses revealed a significant interaction with diagnostic category. Among the 91 participants with microbiologically confirmed TB meningitis, 34% of placebo recipients experienced new infarcts or death by day 60, compared with 15% in the low-dose aspirin group and 11% in the high-dose aspirin group. Does this mean aspirin works for confirmed TB meningitis but not for suspected cases? Maybe. Or maybe the study just wasn't large enough to detect effects in the unconfirmed cases. [Slide 13] That's the challenge with subgroup analyses: they can provide important insights, but they also risk false positives when we slice the data many different ways. Pre-specified subgroup analyses, planned before we look at the data on a biological or theoretical rationale, provide stronger evidence. Post-hoc analyses generate hypotheses that need confirmation in future studies. [Slide 14] This connects directly to our chapter on external validity. The variables that drive heterogeneity within a trial, the effect modifiers, are the same variables that determine whether its findings will travel to new populations. If maternal BMI modifies the effect in our trial, then differences in BMI distributions between the study population and a target population will shape whether our results apply there. In this illustrative forest plot, the diamond at the top is the overall average effect. The subgroup estimates show an intervention that works much better for underweight mothers and when started early in pregnancy, and that benefits female infants more than males. [Slide 15] Even when an RCT demonstrates clear benefit and cost-effectiveness, this is where most interventions fail, because scaling is hard. Pragmatic design choices help close the gap between efficacy and effectiveness: broad eligibility, routine clinical settings, usual care as the comparator, and outcomes that matter to patients and providers. But even a pragmatic effectiveness trial doesn't solve the harder problem of scale-up and dissemination across a large population or health system. [Slide 16] The oral cholera vaccine illustrates this well. Initial efficacy trials in Asia established that it could prevent cholera under controlled trial conditions. The Haiti demonstration project tested effectiveness in an outbreak setting with operational delivery at scale. That evidence led to WHO endorsement of the vaccine for both endemic and epidemic use. But translating the endorsement into widespread deployment required addressing vaccine supply chains, financing mechanisms, targeting strategies, and integration with other cholera control measures, challenges that extend far beyond what any single RCT can address. [Slide 17] Or consider the HIV treatment-as-prevention evidence from HPTN 052, which we saw at the start of this chapter. The trial established that early antiretroviral treatment dramatically reduces HIV transmission in serodiscordant couples. But achieving population-level impact required test-and-treat strategies that could identify HIV-positive people early, link them to care, support adherence, and keep them in treatment. Coverage of antiretroviral therapy during pregnancy in sub-Saharan Africa rose from 33% in 2010 to 69% in 2019, and mother-to-child transmission declined accordingly. But progress has been uneven across the region, reflecting differences in health system capacity, not differences in the underlying evidence. [Slide 18] We started this chapter with a simple mixture of salt, sugar, and water that has saved more than 50 million lives. Randomization creates the counterfactual we can never directly observe, and that's what lets us say causes, and not merely correlates with. But the ORT story also shows what trials can't do on their own. The trials showed that the solution could prevent death from dehydration. They didn't explain how to train millions of mothers to mix it correctly, how to sustain supply chains across rural Bangladesh, or how to overcome physician resistance to something that seemed too simple to work. Getting from a trial result to more than 50 million lives saved required decades of implementation science, community health worker programs, policy advocacy, and political will. [Slide 19] RCTs are our strongest tool for establishing that an intervention causes the outcomes we care about. And they're only the beginning of the work required to turn that evidence into better health.