Chapter 5 · Video 4

Effect sizes, confidence intervals, and how results get misread

8 slides · Video page · All videos · Transcript
Print layout
Slide 1

Effect sizes, confidence intervals, and how results get misread

Slide 2

The authors didn’t report a p-value. They didn’t need to.

The interval excludes 1.0, so the result is “significant” at conventional thresholds
You could derive an approximate p-value from it: roughly 0.02
The CI does everything the p-value does, and more
Slide 3

A p-value tells you how surprising your data would be if there were no effect

Nothing more than that
A “significant” result can be trivially small
A “non-significant” result can reflect an effect the study was too small to detect
Slide 4

An interval carries magnitude and precision at the same time

A narrow interval suggests precise estimation
A wide interval indicates substantial uncertainty
For paper reading: the range of plausible effect sizes given the data
Slide 5
Smith et al., 2023

Specificity of 74.5%, 95% CI 56.9 to 86.6

Visual inspection with acetic acid, in a review of cervical cancer screening in LMICs
Specificity might be as low as 57%, or as high as 87%, depending on context
You can’t act as if it is precisely 74.5%
Slide 6

Two results, and what each one licenses you to say

RR 0.70 (0.55–0.90)
A 30% reduction, best estimate
Plausibly 45% down to 10%
Interval below 1.0: “significant”
RR 0.80 (0.60–1.05)
A 20% reduction, best estimate
Plausibly 40% down to a 5% increase
Interval includes 1.0
Slide 7
Glasgow et al., 2013

Is one kilogram clinically meaningful?

A weight loss intervention reported statistically significant reduction
The absolute difference was 1 kg, 95% CI −2.0 to −0.03
Small effects aren’t automatically dismissible if delivery is cheap at scale
Slide 8
One of the chapter's six

Non-significance is not evidence of no effect

It means you could not rule out chance with the desired confidence
p = 0.049 is not meaningfully different from p = 0.051
The chapter lists six common mistakes in interpretation; the other five are on the page
Next: the same questions asked of tables and figures

Effect sizes, confidence intervals, and how results get misread

Slide 1The result we just took apart came with an interval and no p-value. That absence is worth dwelling on, because it points at what a confidence interval does that a p-value cannot.

The authors didn’t report a p-value. They didn’t need to.

The interval excludes 1.0, so the result is “significant” at conventional thresholds
You could derive an approximate p-value from it: roughly 0.02
The CI does everything the p-value does, and more
Slide 2Notice something important: the authors didn't report a p-value. They didn't need to. Because the confidence interval excludes 1.0 entirely, we know the result would be statistically significant at conventional thresholds. If you're curious, you can derive an approximate p-value from the reported interval; it works out to roughly 0.02 under standard assumptions. But notice how much more the confidence interval tells you. Knowing p equals 0.02 confirms the result is unlikely due to chance if the null hypothesis of no difference is true. Knowing the interval spans 0.49 to 0.95 tells you the effect could be anywhere from modest to substantial. The interval does everything the p-value does, and more.

A p-value tells you how surprising your data would be if there were no effect

Nothing more than that
A “significant” result can be trivially small
A “non-significant” result can reflect an effect the study was too small to detect
Slide 3That study did not report a p-value on this comparison, but many studies do. So what is a p-value? A p-value tells you how surprising your data would be if there were truly no effect. Nothing more. The deeper meaning of p-values, their controversies, and the alternatives are the subject of a later chapter. For now, what matters for critical reading is this: a statistically significant result can be trivially small, and a non-significant result can reflect an important effect that the study was too small to detect. P-values do not tell you how big or important an effect is.

An interval carries magnitude and precision at the same time

A narrow interval suggests precise estimation
A wide interval indicates substantial uncertainty
For paper reading: the range of plausible effect sizes given the data
Slide 4A 95% confidence interval provides a range of values compatible with the data. It tells you about the magnitude of the effect and about the precision of the estimate. A narrow confidence interval suggests precise estimation; a wide interval indicates substantial uncertainty. The formal statistical interpretation is subtler than this, and a later chapter explores it. For practical paper reading, think of the interval as the range of plausible effect sizes given the data.
Smith et al., 2023

Specificity of 74.5%, 95% CI 56.9 to 86.6

Visual inspection with acetic acid, in a review of cervical cancer screening in LMICs
Specificity might be as low as 57%, or as high as 87%, depending on context
You can’t act as if it is precisely 74.5%
Slide 5Here is what a wide one looks like. A systematic review of cervical cancer screening in low- and middle-income countries found that visual inspection with acetic acid had specificity of 74.5%, with a wide 95% confidence interval of 56.9 to 86.6. That wide interval tells you there's substantial heterogeneity across studies. Specificity might be as low as 57% or as high as 87% depending on context. You can't act as if specificity is precisely 74.5% when the data are consistent with values anywhere in that range.

Two results, and what each one licenses you to say

RR 0.70 (0.55–0.90)
A 30% reduction, best estimate
Plausibly 45% down to 10%
Interval below 1.0: “significant”
RR 0.80 (0.60–1.05)
A 20% reduction, best estimate
Plausibly 40% down to a 5% increase
Interval includes 1.0
Slide 6Two worked examples, side by side. The first reports a risk ratio of 0.70, with a 95% interval of 0.55 to 0.90. That means the intervention reduced risk by 30%, with the true effect plausibly ranging from a 45% reduction to a 10% reduction. Because the entire interval is below 1.0, the effect is statistically significant. But note that the range from 10% to 45% is quite wide. You're fairly certain there's some benefit; the magnitude is uncertain. Now consider a risk ratio of 0.80, with an interval of 0.60 to 1.05. The point estimate suggests a 20% reduction, but the interval includes 1.0, so the effect is not statistically significant. Yet look at the range. The data are consistent with a 40% reduction, a 20% reduction, no effect, or even a 5% increase in risk. The non-significant result doesn't tell you the intervention doesn't work. It tells you the study couldn't distinguish between benefit, no effect, and modest harm.
Glasgow et al., 2013

Is one kilogram clinically meaningful?

A weight loss intervention reported statistically significant reduction
The absolute difference was 1 kg, 95% CI −2.0 to −0.03
Small effects aren’t automatically dismissible if delivery is cheap at scale
Slide 7So does size matter? Sometimes researchers just want to know if there's any effect. Some cognitive psychology studies aim to demonstrate that a phenomenon exists at all. But in global health, we usually care about magnitude. A weight loss intervention reported statistically significant weight reduction compared to controls, but the absolute difference was only one kilogram. Is one kilogram clinically meaningful? That's a question p-values can't answer, and the effect size helps you judge. That said, small effects aren't automatically dismissible. An intervention with a modest effect size might still be worth implementing if delivery costs are low and the result is cost-effective at scale. The question isn't just how big, but big enough to justify the investment.
One of the chapter's six

Non-significance is not evidence of no effect

It means you could not rule out chance with the desired confidence
p = 0.049 is not meaningfully different from p = 0.051
The chapter lists six common mistakes in interpretation; the other five are on the page
Next: the same questions asked of tables and figures
Slide 8The chapter lists six common mistakes in interpreting results. Here is the one that does the most damage: interpreting non-significance as evidence of no effect. Not statistically significant means you could not rule out chance with the desired confidence. The effect might still be real, and in underpowered studies it very often is. Closely related is treating 0.05 as a bright line. A p-value of 0.049 is not meaningfully different from a p-value of 0.051, and the arbitrary threshold creates false certainty on both sides of it. The other five mistakes are all worth knowing, and they are waiting for you in the chapter, along with the guidance on reading data visualizations. Head over to the page for those. Underneath all six is the same job: when you read a result, judge importance, and not only significance. Next, the same skeptical reading applied to the tables and figures where most of these numbers actually live.