From a trial result to policy
Slide 1Two trial designs exist because of how programs actually reach people, and once a trial has a result, the work of turning it into policy and practice begins.
When every cluster will eventually receive the program
Withholding the intervention indefinitely is ethically or politically problematic
Limited trainers, equipment, or funding make a phased rollout the only option
Slide 2Standard RCTs keep distinct intervention and control groups throughout the study. But what if it's ethically or politically problematic to withhold the intervention from control clusters indefinitely? Or what if we simply can't turn the intervention on everywhere at once, because we have limited trainers, equipment, or funding, and a phased rollout is the only feasible option?
Stepped-wedge design
Randomizing the order in which clusters start
All clusters start in the control condition
They cross over to the intervention at randomly determined times
By the end, everyone has received the intervention
Slide 3If we're willing to randomize the order in which clusters receive the intervention, a stepped-wedge design becomes possible. All clusters start in the control condition and then cross over to the intervention at randomly determined time points. By the end of the study everyone has received the intervention, but we've still been able to compare intervention and control conditions.
Shah et al., 2015
Active case finding for TB contacts across 34 health centers in Peru
Every center eventually implemented active case finding
Start times were randomly determined over 20 months
Centers that had started were compared with centers that hadn’t
Slide 4The Peru tuberculosis study from our video on designing a trial used this design. All 34 health care centers in the district eventually implemented active case finding for TB contacts, but they started at different, randomly determined times over 20 months. That let the researchers compare TB diagnosis rates in centers that had already begun active case finding with centers that hadn't yet started, while respecting the health department's need to eventually implement the program everywhere.
Parallel cluster designs and the stepped-wedge design
Parallel design (a): clusters randomized once. Stepped wedge (c): clusters cross from control to intervention in sequence; (d) adds transition periods. Source: Hemming et al., 2015, CC BY.
Slide 5In a conventional parallel design, clusters are randomized once and stay in their assigned condition. The stepped-wedge design crosses clusters from control to intervention in sequence, and a variation adds transition periods, time for the intervention to be implemented before data collection resumes. The design works well when a program will be scaled up regardless, so the question is whether it works, when sequential rollout is logistically necessary, and when we can plausibly assume that underlying trends are similar across clusters.
Non-inferiority
Is the new approach good enough?
Superiority: is the intervention better than the comparator?
Non-inferiority: is it not meaningfully worse?
Because it is cheaper, easier to deliver, or has fewer side effects
Slide 6Most trials we'll encounter are superiority trials. They test whether an intervention is better than a comparator. But sometimes the goal is to show that a new treatment is not meaningfully worse, perhaps because it's cheaper, easier to administer, has fewer side effects, or can be manufactured locally. That's called non-inferiority.
Non-inferiority questions in global health
Shorter TB treatment regimens against standard 6-month therapy
Community health worker delivery against facility-based delivery
Lower-cost biosimilars against branded biologics
Heat-stable vaccines against cold-chain-dependent versions
Slide 7Global health is full of non-inferiority questions, like testing shorter TB treatment regimens against standard six-month therapy, or comparing community health worker delivery of a service against facility-based delivery. In each case, the question is whether the new approach is good enough that its practical advantages justify adoption.
A 30% reduction in mortality and a 10-point margin
The standard treatment reduces mortality by 30%
The non-inferiority margin is 10 percentage points
The new treatment must reduce mortality by at least 20%
Slide 8A non-inferiority trial tests whether a new treatment's effect falls within a pre-specified margin of the standard treatment's effect. If the standard treatment reduces mortality by 30%, and we define a non-inferiority margin of 10 percentage points, we're asking: does the new treatment reduce mortality by at least 20%?
Four possible results against the margin
The standard treatment reduces mortality by 30%. With a 10-point margin, the new treatment must reduce mortality by at least 20% to be non-inferior. Reproduced from Chapter 10.
Slide 9This figure shows the four possible outcomes. Everything hinges on where the confidence interval falls relative to two lines: the standard treatment's effect of 30%, and the non-inferiority margin, a reduction of at least 20%. Only when the entire confidence interval stays to the right of the margin can we claim non-inferiority.
Equivalence
Equivalence trials ask a two-sided question
Are two treatments close enough to be interchangeable?
Most common in pharmaceutical regulation, such as generic drugs
Slide 10Non-inferiority is a one-sided question: is the new treatment not too much worse? A related but distinct goal is equivalence. Equivalence trials ask a two-sided question: are two treatments close enough in effect that we can consider them interchangeable? They're most common in pharmaceutical regulation, where a generic drug manufacturer needs to show that its product performs the same as the branded version.
Average effects can mask who benefits
A significant result is where translation into policy and practice begins
Slide 11Now suppose the trial found that the intervention significantly improved the outcome we care about. Translating that finding into policy and practice is one of the hardest parts. Most RCTs report average treatment effects, the average difference in outcomes between intervention and control groups across everyone enrolled. But average effects can mask important heterogeneity in who benefits.
Mai et al., 2018
Aspirin for tuberculous meningitis in Vietnam
All 120 participants: no statistically significant benefit
Planned subgroup, 91 with confirmed TB meningitis: new infarcts or death by day 60
Placebo 34%; aspirin 81 mg 15%; aspirin 1000 mg 11%
Slide 12The tuberculous meningitis trial in Vietnam, which we met in our video on control conditions, illustrates this. The primary analysis across all 120 participants didn't show a statistically significant benefit from aspirin. But planned subgroup analyses revealed a significant interaction with diagnostic category. Among the 91 participants with microbiologically confirmed TB meningitis, 34% of placebo recipients experienced new infarcts or death by day 60, compared with 15% in the low-dose aspirin group and 11% in the high-dose aspirin group. Does this mean aspirin works for confirmed TB meningitis but not for suspected cases? Maybe. Or maybe the study just wasn't large enough to detect effects in the unconfirmed cases.
Pre-specified and post-hoc subgroup analyses
Pre-specified
Planned before looking at the data
Post-hoc
Hypotheses to confirm in future studies
Slide 13That's the challenge with subgroup analyses: they can provide important insights, but they also risk false positives when we slice the data many different ways. Pre-specified subgroup analyses, planned before we look at the data on a biological or theoretical rationale, provide stronger evidence. Post-hoc analyses generate hypotheses that need confirmation in future studies.
An illustrative forest plot of effect modification
Illustrative data, not the Vietnam trial: a schematic forest plot inspired by meta-analyses of micronutrient supplementation trials in LMICs. Reproduced from Chapter 10.
Slide 14This connects directly to our chapter on external validity. The variables that drive heterogeneity within a trial, the effect modifiers, are the same variables that determine whether its findings will travel to new populations. If maternal BMI modifies the effect in our trial, then differences in BMI distributions between the study population and a target population will shape whether our results apply there. In this illustrative forest plot, the diamond at the top is the overall average effect. The subgroup estimates show an intervention that works much better for underweight mothers and when started early in pregnancy, and that benefits female infants more than males.
Scaling up is harder than showing efficacy
Pragmatic designs: broad eligibility, routine settings
Usual care as the comparator; outcomes that matter to patients and providers
Even a pragmatic trial leaves scale-up and dissemination unsolved
Slide 15Even when an RCT demonstrates clear benefit and cost-effectiveness, this is where most interventions fail, because scaling is hard. Pragmatic design choices help close the gap between efficacy and effectiveness: broad eligibility, routine clinical settings, usual care as the comparator, and outcomes that matter to patients and providers. But even a pragmatic effectiveness trial doesn't solve the harder problem of scale-up and dissemination across a large population or health system.
Sévère et al., 2016
Oral cholera vaccine, from efficacy trials to deployment
Efficacy trials in Asia under controlled conditions
Effectiveness in the Haiti outbreak, with delivery at scale
WHO endorsement, then supply chains, financing, targeting, and integration
Slide 16The oral cholera vaccine illustrates this well. Initial efficacy trials in Asia established that it could prevent cholera under controlled trial conditions. The Haiti demonstration project tested effectiveness in an outbreak setting with operational delivery at scale. That evidence led to WHO endorsement of the vaccine for both endemic and epidemic use. But translating the endorsement into widespread deployment required addressing vaccine supply chains, financing mechanisms, targeting strategies, and integration with other cholera control measures, challenges that extend far beyond what any single RCT can address.
Cohen et al., 2012; Astawesegn et al., 2022
HIV treatment as prevention, from HPTN 052 to test-and-treat
Early ART dramatically reduced transmission in serodiscordant couples
ART coverage in pregnancy, sub-Saharan Africa: 33% in 2010, 69% in 2019
Uneven progress, reflecting health system capacity
Slide 17Or consider the HIV treatment-as-prevention evidence from HPTN 052, which we saw at the start of this chapter. The trial established that early antiretroviral treatment dramatically reduces HIV transmission in serodiscordant couples. But achieving population-level impact required test-and-treat strategies that could identify HIV-positive people early, link them to care, support adherence, and keep them in treatment. Coverage of antiretroviral therapy during pregnancy in sub-Saharan Africa rose from 33% in 2010 to 69% in 2019, and mother-to-child transmission declined accordingly. But progress has been uneven across the region, reflecting differences in health system capacity, not differences in the underlying evidence.
What the ORT trials could not do on their own
Train millions of mothers to mix the solution correctly
Sustain supply chains across rural Bangladesh
Overcome physician resistance to something that seemed too simple to work
Slide 18We started this chapter with a simple mixture of salt, sugar, and water that has saved more than 50 million lives. Randomization creates the counterfactual we can never directly observe, and that's what lets us say causes, and not merely correlates with. But the ORT story also shows what trials can't do on their own. The trials showed that the solution could prevent death from dehydration. They didn't explain how to train millions of mothers to mix it correctly, how to sustain supply chains across rural Bangladesh, or how to overcome physician resistance to something that seemed too simple to work. Getting from a trial result to more than 50 million lives saved required decades of implementation science, community health worker programs, policy advocacy, and political will.
In Closing
RCTs are the beginning of the work
Our strongest tool for establishing that an intervention causes the outcomes we care about
And only the beginning of turning that evidence into better health
Slide 19RCTs are our strongest tool for establishing that an intervention causes the outcomes we care about. And they're only the beginning of the work required to turn that evidence into better health.