Running the trial: fidelity, monitoring, and reporting Chapter 10: Randomized Controlled Trials — Video 5 https://ghrbook.com/videos/running-the-trial/ [Slide 1] Once a trial is designed and approved, we actually have to do the study. That means enrolling participants, delivering interventions, collecting data, and monitoring quality, and then reporting what happened. [Slide 2] Sometimes RCTs fail, and the reason isn't that the intervention doesn't work. It's that the intervention was never actually delivered as intended. [Slide 3] This is the challenge of implementation fidelity: making sure the intervention is delivered consistently and as designed to everyone randomized to receive it. In pharmaceutical RCTs, fidelity is relatively straightforward, because we can verify that participants took their pills by measuring blood levels of the drug. In complex behavioral interventions or health system changes, fidelity is much harder to ensure and verify. [Slide 4] A trial in Bangladesh of the Cholera Hospital-based Intervention for 7 Days is a good example of thinking seriously about implementation. The intervention delivered health education through facility-based counseling, mobile health messages, and home visits, and the research team tracked four implementation outcomes. Reach: what proportion of enrolled households actually received each intervention component? Dose: how many mobile health messages were delivered? How many home visits occurred? Fidelity: were intervention messages delivered as intended, covering all key topics? And acceptability: did participants find the intervention acceptable and useful? [Slide 5] By measuring those outcomes systematically, the researchers could distinguish between an intervention that doesn't work, where participants received it as intended and outcomes still didn't improve, and an intervention that wasn't delivered properly. This kind of process evaluation should be embedded in every global health RCT. Without it, a negative result leaves us with profound uncertainty. Did the intervention fail because the underlying theory was wrong, because the intervention was poorly designed, or because implementation challenges kept it from ever really happening? [Slide 6] If we wait until data collection is complete to examine data quality, we've waited too long. But there's an important distinction between monitoring process and monitoring outcomes, and getting it wrong can compromise both our blinding and our statistical validity. The study team can and should review data at least weekly for missing data, implausible values, or signs of fabrication, visit sites, monitor enrollment and follow-up, and track serious adverse events, all without breaking the blind. Looking at outcome data by treatment arm is a different matter entirely. [Slide 7] Examining outcome data is the province of the Data Safety Monitoring Board, or DSMB. It's an independent committee, typically clinicians, biostatisticians, and ethicists who aren't involved in the trial. The DSMB sees unblinded outcome data that the investigators don't, and reviews it at pre-specified interim analyses. It has the authority to recommend stopping the trial early for three reasons: clear benefit has been demonstrated and continuing would be unethical, unexpected harm has emerged, or futility analyses suggest no benefit will be found even with full enrollment. [Slide 8] This raises a statistical subtlety. Every time we look at outcome data and potentially stop the trial, we increase the chance of a false positive. If we test our data five times during a trial at an alpha of 0.05 each time, the overall false positive rate climbs to about 14%, not 5%. [Slide 9] Group sequential methods solve this by adjusting the significance threshold at each interim look. The most common approach, O'Brien-Fleming boundaries, uses very stringent criteria early, when data are sparse. We might need a p-value below 0.0001 to stop. The criteria relax later, and the final analysis uses something close to 0.05. That way the overall Type I error rate stays at 5% across all the looks combined. [Slide 10] The male circumcision trials in Africa are one of the most striking examples of ethical stopping in global health. Three independent RCTs tested whether voluntary medical male circumcision reduced HIV acquisition in men: one in Orange Farm, South Africa, one in Kisumu, Kenya, and one in Rakai, Uganda. [Slide 11] The South African trial was stopped first. At a planned interim analysis in 2005, the DSMB found a 60% reduction in HIV incidence among circumcised men, an effect so large that continuing to randomize men to the control group would have meant knowingly withholding a protective intervention. The Kenya and Uganda trials followed shortly after, stopped by their own DSMBs after showing 53% and 51% reductions. [Slide 12] Stopping can also go the other direction. The HVTN 702 HIV vaccine trial in South Africa was stopped for futility when the independent data monitoring committee found no evidence of efficacy. In both cases, stopping for benefit and stopping for futility, the monitoring boards faced a genuine tension between gathering more data and acting on what they already knew. Group sequential methods exist to make those decisions rigorous. [Slide 13] Setting up a DSMB requires recruiting independent experts willing to serve and establishing a charter that specifies when interim analyses will occur and what criteria will trigger stopping recommendations. For smaller trials or lower-risk interventions, a simpler safety monitoring plan may suffice. But for any trial where serious adverse events are plausible, independent safety monitoring isn't optional. [Slide 14] When it's time to write up the results, we need to report them in a way that lets readers evaluate the quality and validity of our findings. The CONSORT statement, the Consolidated Standards of Reporting Trials, provides an evidence-based minimum set of recommendations for reporting RCTs. A trial might be well designed and carefully conducted, but if the paper doesn't describe the randomization process, show the flow of participants, or present baseline characteristics, readers can't tell whether the results are trustworthy. CONSORT has two core components, the flow diagram and the checklist. [Slide 15] The flow diagram tracks participants through the trial, from initial assessment for eligibility through randomization, allocation, follow-up, and final analysis. This is the flow diagram from the Orange Farm trial we just saw stopped, which randomized 3,274 uncircumcised men. How many people were screened? How many were excluded, and why? [Slide 16] How many were lost to follow-up in each arm? How many were included in the final analysis? This single figure tells readers more about a trial's integrity than pages of prose. When these numbers don't add up, we know something went wrong. [Slide 17] The CONSORT checklist is a list of 25 items covering everything a reader needs to evaluate a trial, and many journals require it at submission. I'd recommend downloading it before you start writing up your results, because it's much easier to collect the right information prospectively than to reconstruct it after the fact. CONSORT also has extensions for trial types common in global health, including cluster randomized trials, non-inferiority trials, and CONSORT-Equity, for health equity trials. [Slide 18] Randomization gives us a fair comparison at baseline. But between randomization and analysis, things go wrong, and that's where we'll pick up in the next video.