Video 3 of 5
Retaining or rejecting the null hypothesis
Once the null distribution exists, the decision about your result is close to mechanical. Alpha as a long-run error rate chosen before the study is run, the p-value as a conditional probability, and false positives and false negatives shown across simulated studies.
6:17 · 14 slides · printable slides · transcript
Slides
Printable deck →▶Transcript14 sections
Generated from the narration script. Plain text version.
1In the last video we built an imaginary world of 10,000 studies in which the treatment did nothing. This video sets the goal posts on that curve, makes the decision the Frequentist approach tells us to make, and then shows you what it looks like when that decision is wrong.
2A skeptical student might say, "I get that a study result does not have to be exactly zero for the null to be true. But how different from zero must a result be for Frequentists to reject the null hypothesis that the treatment had no effect?" Good question. In the Frequentist approach, you decide if the data are extreme relative to what is plausible under the null by setting some goal posts before you conduct the study. If the result crosses the threshold, it's automatically deemed statistically significant and the null hypothesis is rejected. If it falls short, the null hypothesis of no difference is retained. This threshold is known as the alpha level.
3For Frequentists, alpha represents how willing they are to be wrong in the long run when it comes to rejecting the null hypothesis. Traditionally, scientists set alpha to be no greater than 5%. That convention has been questioned. One argument is that the optimal alpha level will sometimes be lower and sometimes be higher than the current convention of 0.05, which they describe as arbitrary. Think before you experiment and justify your alpha.
4Returning to our simulated results, you can see that I drew the goal posts as dotted red lines. They are positioned so that 5% of the distribution of study results falls outside of the lines, 2.5% in each tail. This is a two-tailed test, meaning that we'd look for a result in either direction. The treatment group gets better, on the left, a negative difference. Or worse, on the right. You can also draw a single goal post that contains the full alpha level. That's appropriate when you have a directional alternative hypothesis.
5Step 4 is to make a decision about the null. With the goal posts set, the decision is automatic. Your actual study result either falls inside or outside the goal posts, and you must either retain or reject the null hypothesis. I'll say it again: statistical inference happens on the null.
6Going back to our example, imagine that you actually conducted study number 7,501 of 10,000. When you collected endline data, you found a mean difference between the two arms of negative 3.8. What's your decision with respect to the null? Do you retain, or reject?
7Reject. A difference of negative 3.8 falls outside of the alpha level you set at 5%. Therefore, you automatically reject the null hypothesis and label the result statistically significant. But there's something else we know about this result, the raw effect size of negative 3.8. It falls at the first percentile of our simulated collective. It's an extreme result relative to what we expected if the null is true. We'd say it has a p-value of 1%.
8The p-value is a conditional probability. It's the probability of observing a result as big or bigger than our study result, negative 3.8, and here's the conditional part, if the null hypothesis is true.
9I can hear you muttering to yourself. I think you said, "Why does he keep saying if the null hypothesis is true, like some lawyer who loves fine print?" I do love fine print, but it's important to say this again. When it comes to inference, we never know the truth. We do not know if the null hypothesis is actually true or false. That's why the p-value is a conditional probability. A p-value is the probability of observing the data if the null hypothesis is true, not the probability that the null hypothesis is true given the data. Let that sink in. The p-value might not mean what you want it to mean.
10Furthermore, since we can't know the truth about the null hypothesis, it's possible that we make the wrong decision when we reject or retain it. If the null hypothesis is really true, and that's what I simulated, we made a mistake by rejecting the null. We called the result statistically significant, but this is a false positive. Statisticians refer to this mistake as a Type I error, though I think the term false positive is more intuitive, since we're falsely claiming that our treatment had an effect when it did not. The good news is that in the long run we will only make this mistake 5% of the time if we stick to the Frequentist approach. The bad news is that we have no way of knowing if this experiment is one of the times we got it wrong.
11The other type of mistake we can make is called a Type II error, a false negative. We make this mistake when the treatment really does have an effect, but we fail to reject the null. We'll talk more about false negatives, and power, when we get to sample size.
12Why did I bother simulating 10,000 studies instead of just showing you the formulas? Because simulation makes the abstract concrete. When you watch studies pile up into a bell curve, the Central Limit Theorem stops being a theorem and starts being something you can see. When you watch 20 studies flash by and only 1 rejects the null, the 5% false positive rate becomes visible. Here are 20 of the 10,000 studies I simulated, all with no true treatment effect. Study number 7,501, boxed at the bottom left, is the only one in this set that crosses the goal post and rejects the null. In a world where the treatment doesn't work, this study would lead us to falsely conclude that it does. That's a Type I error.
13So what did the real HAP trial find? After adjusting for the study site and participants' depression scores at baseline, the program reduced depression severity by an average of 7.57 points, with a 95% confidence interval of negative 10.27 to negative 4.86. Look back at that null distribution and you'll see that a difference of this size is off the chart. You would not expect to find an effect size this big if the null hypothesis of no difference was true.
14The 95% confidence interval excludes the null effect of 0, so we know the p-value will be less than 5%. And it was. The reported p-value was less than 0.0001. If the null hypothesis is true, meaning that there really was no treatment effect, you would expect to get a difference of 7.5 or more less than 0.01% of the time. In the next video we look at what a p-value does not tell you, and at why so many people want to get rid of it.