After randomization: attrition, non-compliance, and spillover Chapter 10: Randomized Controlled Trials — Video 6 https://ghrbook.com/videos/after-randomization/ [Slide 1] Randomization gives us a fair comparison at baseline. Between randomization and analysis, three things can erode that comparison: participants drop out, participants don't follow the plan, and the treatment given to some people affects others. [Slide 2] Participants drop out. Others don't take the treatment they were assigned. These threats don't invalidate the RCT, but ignoring them can. We'll start with attrition, then non-compliance, and then spillover, which is the reason we randomize clusters. [Slide 3] Some loss to follow-up is inevitable in any trial. Participants move, lose interest, get sick, or die. The immediate cost is statistical. Fewer participants means less power to detect an effect. If we planned for 200 per arm and 30% drop out, we're now running an underpowered study. [Slide 4] But the real danger isn't the lost power. It's bias, when attrition is systematic. If sicker participants in the intervention arm drop out because the treatment has side effects, while healthier participants stay, the remaining sample no longer represents the population we randomized. The treatment group looks artificially healthy, and the comparison is no longer fair. [Slide 5] How do we assess whether attrition is a problem? We start by comparing the baseline characteristics of completers and dropouts in each arm. If dropouts look similar to completers on measured characteristics, and attrition rates are similar across arms, the threat is lower. If dropouts in the treatment arm look systematically different from dropouts in the control arm, we have a problem that no statistical method can fully solve. [Slide 6] When attrition does occur, sensitivity analysis is our main tool. The most conservative approach is to assume the worst: everyone who dropped out of the treatment arm had bad outcomes, and everyone who dropped out of the control arm had good outcomes. If the results hold under that extreme assumption, they're robust. More commonly, we'll use multiple imputation or inverse probability weighting to account for missing data under various assumptions, and report how sensitive our conclusions are to those assumptions. [Slide 7] Non-compliance takes two forms. Participants assigned to the intervention may not take it up. They skip sessions, don't take the medication, or never show up at all, since participation is voluntary. And participants assigned to control may obtain the intervention on their own. Both dilute the contrast between arms. [Slide 8] How we analyze the data in the face of non-compliance depends on the question we want to answer. There are three analyses, and each one answers a different question. [Slide 9] Intention-to-treat analysis keeps participants in their randomized group regardless of what they actually did. Someone randomized to the intervention who never showed up? Still analyzed as intervention. That preserves the benefits of randomization, and it answers a policy-relevant question: what happens when we offer this intervention to a population, knowing that real-world uptake will be imperfect? It's conservative, but intention-to-treat is the standard primary analysis for RCTs and should always be reported. [Slide 10] Per-protocol analysis restricts to participants who actually received the intervention as intended, and control participants who didn't cross over. It answers a different question: what is the effect in people who actually adhere? The problem is that adherers differ systematically from non-adherers. They tend to be healthier, more motivated, and to have better social support. So per-protocol analysis can introduce selection bias, even though it seems to answer a more clinically relevant question. [Slide 11] Complier average causal effect analysis, sometimes called the treatment effect on the treated, uses instrumental variables to estimate the effect of actually receiving the intervention among people who would comply if offered it. Unlike per-protocol analysis, it accounts for the fact that compliance isn't random. It uses randomization itself as the instrument: being assigned to the treatment arm only affects your outcome through receiving treatment, if you're a complier. It recovers a causal estimate, but only for the subpopulation of compliers. Most RCTs should report intention-to-treat as the primary analysis, with per-protocol and, where relevant, the complier average causal effect as supplementary analyses. Together they tell a more complete story than any single approach. [Slide 12] Spillover is one of the most underappreciated complications in trial design. It occurs when the treatment given to one person affects the outcomes of another person who wasn't treated. From a public health perspective, spillover is often the whole point. We want deworming one child to reduce transmission to nearby children. We want vaccinating enough people to create herd immunity. We want health information shared with one household to spread through a community. [Slide 13] The classic example is a study of school-based deworming in Kenya, which found that untreated children in treated schools, and even children in nearby untreated schools, experienced health gains from reduced transmission. [Slide 14] Spillover complicates research because it violates the stable unit treatment value assumption, or SUTVA: the assumption that one person's treatment status doesn't affect another person's outcome. Spillover can be biological, informational, or behavioral. If the control group is also benefiting from the intervention, it isn't really untreated, and our estimated treatment effect shrinks because the comparison group is also benefiting. [Slide 15] This is why cluster randomization helps. We randomize groups when the intervention is delivered at the group level, when spillover between individuals is likely, or when individual randomization would be logistically impossible or culturally inappropriate. By randomizing entire villages or schools, we create a buffer between treatment and control groups. Spillover still happens within clusters, but it's less likely to cross between clusters, especially when they're geographically separated. We can also design studies to measure spillover directly, by comparing untreated people who are near treated clusters with those who are far from them. [Slide 16] Cluster randomization solves the spillover problem, but it introduces a statistical one. People within the same cluster tend to be more similar to each other than to people in different clusters. That similarity is captured by the intracluster correlation coefficient, or ICC, and it means that adding more people from the same cluster gives us less new information than adding people from different clusters. To build intuition, imagine every person in a cluster is an identical clone. Once we've measured one clone's response to treatment, measuring a second tells us nothing new, so the effective sample size is 1. That's an ICC of 1. If people within clusters are no more similar to each other than to people in other clusters, the ICC is 0, and clustering doesn't cost us anything. Reality falls somewhere in between, and even modest ICCs can dramatically inflate the sample size we need. [Slide 17] Suppose we're testing a school-based nutrition program and the outcome is child height. If the ICC is 0.05, meaning 5% of the variation in height is between schools, and there's an average of 50 children per school, the design effect is 3.45. We'd need 3.45 times as many participants in a cluster trial as in an individually randomized trial to achieve the same statistical power. [Slide 18] This figure shows how quickly that adds up. With the same total sample size, a cluster trial can detect much less than an individual trial. [Slide 19] This is why cluster trials require many more participants than individual trials, and why we need to plan for clustering during the sample size calculation. Ignoring clustering leads to underpowered studies and inflated Type I error rates. In the next video, we'll see a cluster design in which every cluster eventually receives the intervention.