Video 1 of 5
What is statistical inference?
A large trial found a 30% lower risk of cancer among older women taking vitamin D3 and calcium, and the authors still reported the result as inconclusive. Working through that disagreement introduces statistical conclusion validity and the two major approaches to inference.
7:16 · 18 slides · printable slides · transcript
Slides
Printable deck →▶Transcript18 sections
Generated from the narration script. Plain text version.
1This chapter is about statistical inference. In an earlier chapter you learned to interpret effect sizes and confidence intervals when critically reading papers. Now we'll go deeper. Why does statistical inference work the way it does? What are we actually claiming when we calculate a p-value or a confidence interval? And why is all of this so controversial? We'll start with a trial that reported a 30% reduction in cancer risk and concluded that supplements do not lower cancer risk.
2Here's a figure from a randomized clinical trial of 2,303 healthy postmenopausal women. The trial set out to answer whether dietary supplementation with vitamin D and calcium reduces the risk of cancer among older women. Before we go any further, look at the image and decide what you think.
3If you said yes, you're probably in good company. I think most readers will come to the same conclusion. Without knowing anything else about the specific analysis, or about statistics in general, you can look at this figure and see that both groups started at 0% of participants with cancer, which makes sense given the design. Over time, members of both groups developed some type of cancer. And by the end of the study period, cancer was more common among the non-supplement group.
4But yes is not what the authors concluded. Here's what they said. Supplementation compared with placebo did not result in a significantly lower risk of all-type cancer at 4 years.
5Technically they are correct. The supplement group had a 30% lower risk for cancer compared to the placebo group, a hazard ratio of 0.70. But the 95% confidence interval around this estimate spanned from 0.47, a 53% reduction, to 1.02, a 2% increase. It crossed the line of no effect for ratios, at 1.0. The p-value was 0.06 and their a priori significance cutoff was 0.05, so the result was deemed not significant, and the conclusion was that supplements do not lower cancer risk.
6But is that the best take? Not everyone thought so. Here's what Ken Rothman and his colleagues wrote to the journal editors.
7The journal ultimately published a different letter to the editor that raised similar issues, and the authors of the original paper responded. They wrote that the possibility that the results were clinically significant should be considered, and that the 30% reduction in the hazard ratio suggests that this difference may be clinically important.
8The best answer, at least in my view, is that the trial was inconclusive. The point estimate is that supplements reduced cancer risk by 30%, but the data are also consistent with a relative reduction of 53%, and with an increase of 2%. In absolute terms, the group difference in cancer prevalence at Year 4 was 1.69 percentage points. It seems like there might be a small effect. Whether a small effect is clinically meaningful is for the clinical experts on your research team to decide.
9This example highlights some of the challenges with statistical inference. Science is all about inference: using limited data to make conclusions about the world. We're interested in this sample of 2,300 because we think the results can tell us something about cancer risk in older women more generally. But to make this leap, we have to make several inferences.
10First, we have to decide whether we think the observed group differences in cancer risk in our limited study sample reflect a true difference. This is a question about statistical inference. Second, we have to ask whether this difference in observed cancer risk was caused by the supplements. This is a question of causal inference, and internal validity. Finally, if we think the effect is real, meaningful, and caused by the intervention, do we think the results apply to other groups of older women? This is a question of generalizability and external validity. Those are the next two chapters.
11This first question, whether we can trust the statistical relationship we've found, has a name. It's a question of statistical conclusion validity. Before we can ask what a relationship means, or where it applies, we have to ask whether there's a relationship at all, and whether we've correctly identified it. Statistical conclusion validity is threatened when we make errors in inference: concluding there's an effect when there isn't, missing a real effect, or misestimating the size of an effect.
12There are two main approaches to statistical inference: the Frequentist approach and the Bayesian approach. A key distinction between the two is the assumed meaning of probability. Believe it or not, smart people continue to argue about the definition of probability. If you are a Frequentist, then you believe that probability is an objective, long-run relative frequency. Your goal when it comes to inference is to limit how often you will be wrong in the long run. If you are a Bayesian, you favor a subjective view of probability that says you should start with your degree of belief in a hypothesis, and update that belief based on the data you collect.
13Before we get too far along, please think about what you want to know most. Is it the probability of observing the data you collected if your preferred hypothesis was not true? Or is it the probability of your hypothesis being true, based on the data you observed?
14Open just about any medical or public health journal and you'll find loads of tables with p-values and asterisks, and results described as significant or non-significant. These are artifacts of the Frequentist approach, specifically the Neyman-Pearson approach. To explore how Frequentist inference works, we'll use a real trial that we'll return to throughout this book: the Healthy Activity Program trial.
15Here's a real problem. Depression is treatable with psychological therapies, but these treatments require trained professionals, like psychiatrists, psychologists, and clinical social workers, who are scarce in most of the world. In India, there are fewer than 1 psychiatrist per 100,000 people. The result is a massive treatment gap: most people with depression receive no effective care.
16What if we could train lay counselors, people without formal mental health credentials, to deliver a simplified but effective psychological treatment? The Healthy Activity Program does exactly this. It is based on behavioral activation, a therapeutic approach that helps people re-engage with meaningful activities and break the cycle of withdrawal and low mood. Lay counselors deliver 6 to 8 sessions of 30 to 40 minutes each.
17The trial enrolled 495 adults with moderately severe to severe depression from 10 primary health centers in Goa, India. Participants were randomly assigned to receive either enhanced usual care alone, the control group, or enhanced usual care plus the program, the treatment group. Enhanced usual care meant that physicians received the patient's screening results and clinical guidelines for treating depression. That's more than typical care, but no structured psychological treatment. The question: does adding the program to enhanced usual care reduce depression severity compared to enhanced usual care alone?
18This is the question we'll use to learn how statistical inference works. In the Neyman-Pearson approach, you set some ground rules for inference, collect and analyze your data, and compare your result to the benchmarks you set. Inference is essentially automatic once you set the ground rules. The next video takes the first two of those steps, and it asks you to imagine 10,000 studies in which the treatment never worked.