Effect sizes, confidence intervals, and how results get misread Chapter 5: How to Read Scientific Articles — Video 4 https://ghrbook.com/videos/effect-sizes-and-confidence-intervals/ [Slide 1] The result we just took apart came with an interval and no p-value. That absence is worth dwelling on, because it points at what a confidence interval does that a p-value cannot. [Slide 2] Notice something important: the authors didn't report a p-value. They didn't need to. Because the confidence interval excludes 1.0 entirely, we know the result would be statistically significant at conventional thresholds. If you're curious, you can derive an approximate p-value from the reported interval; it works out to roughly 0.02 under standard assumptions. But notice how much more the confidence interval tells you. Knowing p equals 0.02 confirms the result is unlikely due to chance if the null hypothesis of no difference is true. Knowing the interval spans 0.49 to 0.95 tells you the effect could be anywhere from modest to substantial. The interval does everything the p-value does, and more. [Slide 3] That study did not report a p-value on this comparison, but many studies do. So what is a p-value? A p-value tells you how surprising your data would be if there were truly no effect. Nothing more. The deeper meaning of p-values, their controversies, and the alternatives are the subject of a later chapter. For now, what matters for critical reading is this: a statistically significant result can be trivially small, and a non-significant result can reflect an important effect that the study was too small to detect. P-values do not tell you how big or important an effect is. [Slide 4] A 95% confidence interval provides a range of values compatible with the data. It tells you about the magnitude of the effect and about the precision of the estimate. A narrow confidence interval suggests precise estimation; a wide interval indicates substantial uncertainty. The formal statistical interpretation is subtler than this, and a later chapter explores it. For practical paper reading, think of the interval as the range of plausible effect sizes given the data. [Slide 5] Here is what a wide one looks like. A systematic review of cervical cancer screening in low- and middle-income countries found that visual inspection with acetic acid had specificity of 74.5%, with a wide 95% confidence interval of 56.9 to 86.6. That wide interval tells you there's substantial heterogeneity across studies. Specificity might be as low as 57% or as high as 87% depending on context. You can't act as if specificity is precisely 74.5% when the data are consistent with values anywhere in that range. [Slide 6] Two worked examples, side by side. The first reports a risk ratio of 0.70, with a 95% interval of 0.55 to 0.90. That means the intervention reduced risk by 30%, with the true effect plausibly ranging from a 45% reduction to a 10% reduction. Because the entire interval is below 1.0, the effect is statistically significant. But note that the range from 10% to 45% is quite wide. You're fairly certain there's some benefit; the magnitude is uncertain. Now consider a risk ratio of 0.80, with an interval of 0.60 to 1.05. The point estimate suggests a 20% reduction, but the interval includes 1.0, so the effect is not statistically significant. Yet look at the range. The data are consistent with a 40% reduction, a 20% reduction, no effect, or even a 5% increase in risk. The non-significant result doesn't tell you the intervention doesn't work. It tells you the study couldn't distinguish between benefit, no effect, and modest harm. [Slide 7] So does size matter? Sometimes researchers just want to know if there's any effect. Some cognitive psychology studies aim to demonstrate that a phenomenon exists at all. But in global health, we usually care about magnitude. A weight loss intervention reported statistically significant weight reduction compared to controls, but the absolute difference was only one kilogram. Is one kilogram clinically meaningful? That's a question p-values can't answer, and the effect size helps you judge. That said, small effects aren't automatically dismissible. An intervention with a modest effect size might still be worth implementing if delivery costs are low and the result is cost-effective at scale. The question isn't just how big, but big enough to justify the investment. [Slide 8] The chapter lists six common mistakes in interpreting results. Here is the one that does the most damage: interpreting non-significance as evidence of no effect. Not statistically significant means you could not rule out chance with the desired confidence. The effect might still be real, and in underpowered studies it very often is. Closely related is treating 0.05 as a bright line. A p-value of 0.049 is not meaningfully different from a p-value of 0.051, and the arbitrary threshold creates false certainty on both sides of it. The other five mistakes are all worth knowing, and they are waiting for you in the chapter, along with the guidance on reading data visualizations. Head over to the page for those. Underneath all six is the same job: when you read a result, judge importance, and not only significance. Next, the same skeptical reading applied to the tables and figures where most of these numbers actually live.