Power Analysis Basics
ARS Jonesboro, July 16, 2026
Why do we do power analysis?
- “Because we have to do it to get our study approved”
- Because it helps us design the study, foreseeing any issues that might arise
- Because it reduces the chance of doing weak and inconclusive studies that are doomed before they even start
- Because it elevates our science, and our science makes the world a better place
At the end of this talk, you will understand …
- What statistical power is and why it’s important
- Why we have to trade off between false positives and false negatives
- Why power depends on sample size
- What effect size is and why statistical power depends on it
- That doing a power analysis after you’ve collected the data is pointless
- That the quality and realism of a power analysis is better, the more work you put into it
A framework for statistical power
- Classical statistical framework: use evidence from the data to try to reject a null hypothesis
- We can never be 100% certain of the truth
- We might be wrong because of natural variation, measurement error, biases in our sample, or any/all of those
Right and wrong in different ways
- Simplest case: binary (yes or no) outcome
- You can be right or wrong in different ways, depending on what the truth is
- False positive = Type I error
- False negative = Type II error
![Two ways to be wrong]()
- \(Y\) is the truth, \(\hat{Y}\) is the model’s prediction
![]()
POWER: If an effect exists, the chance that our study will find it
- Power is important for both experimental and observational studies!
Definition of statistical power
- Power of a statistical test: the probability that the test will detect a phenomenon if the phenomenon is true
- Power is the probability of declaring a true positive if the null hypothesis is false
- In contrast, the p-value is the probability of declaring a false positive if the null hypothesis is true
You can’t be right all the time
- It is impossible to completely eliminate all false positives and all false negatives
- The only way to be 100% certain you will never get a false positive is for your test to give 100% negative results
- Find the sweet spot that reduces both false positives and false negatives to an acceptably low rate
Magic number: false positive rate
- Traditional paradigm of Western science is conservative
- Low probability of false positives, at the cost of a fairly high rate of false negatives
- \(p < 0.05\) comes from significance level \(\alpha = 0.05\): 5% probability of a false positive
Magic number 2: false negative rate
- Power is 1 - false negative rate
- Commonly we target 20% false negative rate, or \(\beta = 0.20\)
- For example IACUC guidelines require 80% power
- \(\frac{0.20}{0.05}\) = 4:1 ratio of probability of a false negative to probability of a false positive
- No reason why we must use a 4:1 ratio
What’s wrong with an underpowered study: false negatives
- An underpowered study is one that has a low probability of detecting a significant effect, if the effect truly exists in the world
- Imagine we do an underpowered study and get a negative result: we can’t reject the null hypothesis
- Because the power is low, we don’t know if we got a true negative or a false negative
- No new knowledge was gained, we might as well not have done the study at all
What’s wrong with an underpowered study: false positives
- But what if the underpowered study gives us a positive (significant) result?
- Some people say that means the study was powerful enough after all, and power only matters if we don’t get a significant result
False positives in underpowered studies, continued
- At small sample sizes when power is low, natural variation/measurement error/sampling bias can cause overestimates of the effect size
- At low power, even though false positive rate is fixed at 0.05, the ratio of false positives:true positives increases
- A single positive result is not very convincing if it comes from an underpowered study
- Publication bias in favor of positive results from small studies makes the problem worse
Err on the side of caution
- If you collect too many samples: excess resources wasted/subjects harmed, but you still learn something
- If you collect the “just right” number: resources used most efficiently to gain knowledge … but we can never know this exact number in advance!
- If you collect too few samples: study is inconclusive and we do not learn anything. All resources spent are wasted