Power Analysis Basics

ARS Jonesboro, July 16, 2026

Why do we do power analysis?

  • “Because we have to do it to get our study approved”
  • Because it helps us design the study, foreseeing any issues that might arise
  • Because it reduces the chance of doing weak and inconclusive studies that are doomed before they even start
  • Because it elevates our science, and our science makes the world a better place

At the end of this talk, you will understand …

  • What statistical power is and why it’s important
  • Why we have to trade off between false positives and false negatives
  • Why power depends on sample size
  • What effect size is and why statistical power depends on it
  • That doing a power analysis after you’ve collected the data is pointless
  • That the quality and realism of a power analysis is better, the more work you put into it

Power basics

A framework for statistical power

  • Classical statistical framework: use evidence from the data to try to reject a null hypothesis
  • We can never be 100% certain of the truth
  • We might be wrong because of natural variation, measurement error, biases in our sample, or any/all of those

Right and wrong in different ways

  • Simplest case: binary (yes or no) outcome
  • You can be right or wrong in different ways, depending on what the truth is
    • False positive = Type I error
    • False negative = Type II error

Two ways to be wrong

  • \(Y\) is the truth, \(\hat{Y}\) is the model’s prediction

POWER: If an effect exists, the chance that our study will find it

  • Power is important for both experimental and observational studies!

Definition of statistical power

  • Power of a statistical test: the probability that the test will detect a phenomenon if the phenomenon is true
  • Power is the probability of declaring a true positive if the null hypothesis is false
  • In contrast, the p-value is the probability of declaring a false positive if the null hypothesis is true

You can’t be right all the time

  • It is impossible to completely eliminate all false positives and all false negatives
  • The only way to be 100% certain you will never get a false positive is for your test to give 100% negative results
  • Find the sweet spot that reduces both false positives and false negatives to an acceptably low rate

Magic number: false positive rate

  • Traditional paradigm of Western science is conservative
  • Low probability of false positives, at the cost of a fairly high rate of false negatives
  • \(p < 0.05\) comes from significance level \(\alpha = 0.05\): 5% probability of a false positive

Magic number 2: false negative rate

  • Power is 1 - false negative rate
  • Commonly we target 20% false negative rate, or \(\beta = 0.20\)
    • For example IACUC guidelines require 80% power
  • \(\frac{0.20}{0.05}\) = 4:1 ratio of probability of a false negative to probability of a false positive
  • No reason why we must use a 4:1 ratio

  • Tradition!

What’s wrong with an underpowered study: false negatives

  • An underpowered study is one that has a low probability of detecting a significant effect, if the effect truly exists in the world
  • Imagine we do an underpowered study and get a negative result: we can’t reject the null hypothesis
  • Because the power is low, we don’t know if we got a true negative or a false negative
  • No new knowledge was gained, we might as well not have done the study at all

What’s wrong with an underpowered study: false positives

  • But what if the underpowered study gives us a positive (significant) result?
  • Some people say that means the study was powerful enough after all, and power only matters if we don’t get a significant result

False positives in underpowered studies, continued

  • At small sample sizes when power is low, natural variation/measurement error/sampling bias can cause overestimates of the effect size
  • At low power, even though false positive rate is fixed at 0.05, the ratio of false positives:true positives increases
  • A single positive result is not very convincing if it comes from an underpowered study
  • Publication bias in favor of positive results from small studies makes the problem worse

Power analysis is a ballpark figure

  • If you knew exactly how large the effect is in your system, you would know exactly how many replicates you need
  • But if you knew exactly how large the effect is in your system, you wouldn’t need to do the study!
  • Power calculations are at best rough estimates

Image (c) Britannica

Err on the side of caution

  • If you collect too many samples: excess resources wasted/subjects harmed, but you still learn something
  • If you collect the “just right” number: resources used most efficiently to gain knowledge … but we can never know this exact number in advance!
  • If you collect too few samples: study is inconclusive and we do not learn anything. All resources spent are wasted

Image (c) Kidsotic

What do we need to know to do a power analysis?

  • Significance level
  • Power
  • Sample size
  • Effect size

(Un)known quantity 1: Significance level

  • Many stat analyses in ag science assume \(\alpha = 0.05\)
  • But \(\alpha\) may need to be adjusted if you have multiple comparisons
  • This adjustment is often (wrongly) neglected when doing power analysis

(Un)known quantity 2: Power

  • Sometimes power analysis is done assuming a desired level of power (often 80%, or \(\beta = 0.20\))
  • Or we may estimate the power to detect a given effect at a given sample size

(Un)known quantity 3: Sample size

  • We may know how many samples/replicates we can (afford to) collect or measure
  • Or we may calculate the minimum sample size needed to get a certain power
  • Often there is a range of feasible sample sizes, not just a single number

Sample size: the more the better, up to a point

  • Statistical inference: take a sample, calculate statistics, use statistics to estimate parameters (true values) of the sampled population
  • The bigger the sample, the closer our statistical estimates get to the true parameters

  • Accuracy increases more slowly as the estimate asymptotically approaches the true value
  • We get diminishing returns of adding more data
  • Where does the curve bend?

(Un)known quantity 4: Effect size

  • The size of a practically meaningful effect that you want to be able to detect statistically
  • Often not a single number but a range
  • Probably the most important quantity for power analysis
  • Also the one that takes the most effort to determine

Statistical power increases as effect size increases

  • The stronger the effect, the easier it is to see that effect above the background noise
  • True positives become more likely and false negatives less likely
  • It’s easy to get high power at low false positive rate if the effect size is large
  • If the effect size is small, our only hope is to collect more data

Standardized effect sizes

  • Effect sizes are almost always in standardized units
  • Different studies measure response variables in different units (biomass in kg or g or lb, height in cm, m, or in)
  • If we convert to standard deviation units, we can compare results that are on different scales

What are some commonly used effect sizes?

Effect Effect size metric Description Example
Difference of the mean of two groups Cohen’s d The difference between the two means divided by their combined standard deviation. How many standard deviations greater is one mean than the other? The effect of a management intervention on biomass yield
Difference in a binary outcome between two groups Odds ratio The ratio of the odds of the outcome happening in one group, to the odds of it happening in another group The effect of a seed treatment on seed germination rate
Strength of relationship between two continuous variables Pearson correlation coefficient r You should know this one! The effect of temperature on soil respiration
Variance explained by treatment group when comparing three or more groups Cohen’s f (Square root of) the ratio of variance explained by group, to the residual variance not explained by group The effect of multiple genotypes and environments on a trait

Estimating effect sizes

How to get information about effect size

  • Data from your own lab
  • The literature
  • Prior external information
  • Effect size benchmarks

Effect sizes from data

  • Find datasets that are roughly comparable to what you expect in your proposed study
    • Datasets where your response of interest is measured in a similar system
    • Datasets where a different variable is measured in your study system
  • Calculate effect sizes from the prior datasets, and design your study around that effect size

Effect sizes from the literature

  • Find studies on a similar topic or done in a similar system
  • Even without their raw data, you can still calculate effect size from summary statistics, tables, or figures
  • Essentially a quick and informal meta-analysis
  • Quality/relevance of the studies is more important than quantity

Example: impact of transgenic Bt crops on abundance of soil organisms

  • We want to measure this well-studied effect in a new system or context
  • We can use previous studies on the same or similar response to get an idea of the range of effect sizes we can expect
  • Design our study to detect an effect on the low end of that range

Image (c) NCSU Entomology

Examples from the literature

  • Find published estimates of soil organismal abundance compared between transgenic Bt line and isogenic line
  • I selected these haphazardly from a 2022 meta-analysis of the topic
  • Assume the Bt effect size in our system will be in the same neighborhood as the previous studies

Frouz et al. 2008, data from Table 1

Cerevkova & Cagan 2015, data from Table 2

Arias-Martins et al. 2016, data from Table 4

Distribution of Cohen’s d from the published tables

What effect size to use?

  • There is no single correct answer; it depends what you can justify
  • We could use the median (\(d = 0.46\), a medium effect size), or a lower quantile
  • Or you could use this distribution to get an idea of the power for a predetermined design
    • For example, if your study has 80% power to detect \(d \geq 0.3\), you’d be happy
    • If your study has 80% power to detect \(d \geq 0.6\), back to the drawing board!

Effect sizes based on external criteria

  • You are a domain-knowledge expert so you can justify effect sizes without any prior data
  • What effect size is biologically/practically/economically relevant?
  • You can consult with stakeholders to determine this
    • Example: a stakeholder is interested in a treatment that will increase yield by 2 bushels/acre, or decrease mortality by 25%
  • The problem is often coming up with a way to convert the stakeholder’s wishes to a standardized metric

Effect size based on no prior information

  • I often hear “Nothing like this has ever been done before, so I have no idea what the effect size should be”
  • Usually this indicates a failure to think creatively
  • You don’t need exactly the same study, anything with a similar response variable, or a different variable in a similar system, will do

Conventional effect size benchmarks

  • If you truly have no reasonable way of estimating effect size, we still have to assume one to calculate power
  • Most standardized effect size metrics have conventional benchmarks or thresholds
d f Effect size
0.20 0.10 Small
0.50 0.25 Moderate
0.80 0.40 Large

Happy medium

  • Common practice: target medium effect size
  • Designing a study to detect a small or weak effect requires a huge sample size
  • Designing a study to detect large effects is not worthwhile
  • I do not recommend blindly using the medium effect size all the time
  • The benchmarks are very general across fields and may not be appropriate for your study system
  • No substitute for finding relevant data or consulting experts/stakeholders

Power analysis using formulas

Plug-and-chug power analysis

  • Certain simple experimental designs have analytical equations for statistical power, with these variables:
    • Sample size
    • Desired false positive rate (alpha, often set to 0.05)
    • Effect size
    • Statistical power
  • Plug all but one into our formula and solve for the unknown

Sample size given power and effect size

  • We are comparing four different groups in our study
  • We want to have 80% power to detect a moderate effect size f = 0.25
  • How many samples do we need?
  • Formula says n = 45 per group

Effect size given power and sample size

  • We can only afford 60 total samples (15 per group)
  • What effect size can we detect at 80% power?
    • minimum detectable effect size (MDES)
  • Formula says we can detect f = 0.44

Power given effect size and sample size

  • Our preliminary analysis tells us we can expect effects around f = 0.4
  • We can afford n = 60 samples
  • What is the power of the study?
  • Formula says power is 71%

Power analysis using simulation

Why do we need to do power analysis by simulation?

  • Plug-and-chug power analysis is the exception, not the rule
  • Most experimental and observational study designs in ag science require mixed models to analyze
  • There is not a simple equation for power
  • Simulation to the rescue!

Basic power simulation workflow

  1. Generate a dataset using a given set of assumptions (effect size, sample size, variance components, etc.)
  2. Fit a statistical model to the dataset and calculate a p-value for your hypothesis of interest
  3. Repeat n times
  4. Your power is the proportion of times, out of n, that you got a significant p-value

A simple simulation example: t-test

  • Assume the following:
    • Treatment mean = 12
    • Control mean = 9
    • Standard deviation of both groups = 3
    • Sample size of both groups = 30
    • Individuals are normally distributed around group mean
  • Simulate 10000 datasets and do a t-test on each one
  • Proportion of the time that \(p < 0.05\) is the power!

Simulation gives same result as formula

set.seed(1)
p_vals <- rep(0, 10000)

for (i in seq_along(p_vals)) {
  x_t <- rnorm(n = 30, mean = 12, sd = 3)
  x_c <- rnorm(n = 30, mean = 9, sd = 3)
  p_vals[i] <- t.test(x_t, x_c)$p.value
}

mean(p_vals < 0.05)
[1] 0.9682
pwr.t.test(n = 30, d = (12 - 9)/3, sig.level = 0.05)$power
[1] 0.9677083

Power by simulation for mixed models

  • Mixed models: linear models that include random effects to account for grouping in the data
  • Also called multilevel models because they have more than one level of variance
    • Example: Individuals nested within blocks nested within environments

Mixed models have more unknowns

  • In addition to effect size and sample size, you also need variance components for each level
  • Statistical power depends on the variance at all levels
  • Split-split-plot or repeated measures in time may have many variance components

Power simulation

  • How many replications should we include in a study design?
  • Here we will simulate random effects for new datasets with different numbers of blocks (reps)
  • Repeat many times for each number of blocks
  • Do a statistical test of choice on each simulated dataset
  • Calculate power as % of times the test rejects the null hypothesis

Power curve from simulation

Concluding remarks

With power analysis, more effort = better results

  • You can do a power analysis just fine by making up plausible numbers
  • But the more work you put into it, the higher the quality of the power analysis … and the resulting science
    • Background research in the literature and your own previous datasets
    • Simulation based on plausible assumptions
  • If you say “Do a power analysis for me” without providing the context, it may not be a good quality analysis

Power analysis is iterative

What you think power analysis is

What power analysis actually is

Credit to Jessica Logan for the images

Type S and Type M error: another way to be wrong

Gelman, Skardhamar, & Aaltonen 2017, J. Quant. Criminology
  • Beyond the “false positive vs. false negative” paradigm (see work of Andrew Gelman and others)
    • Type S: The sign is wrong
    • Type M: The magnitude is wrong
  • Even more impetus to design appropriately powered studies!

Post hoc power analysis is pointless

  • Reviewers in applied science journals often ask scientists to do a power analysis on a study after the fact
  • But if you do a power analysis assuming the observed difference is the true difference, you will always find low power if you didn’t detect, and high power if you did detect
  • If the sample size is low, we don’t know whether the observed difference is the true difference
  • You already knew if the study was underpowered before starting the study
  • The power of a study after the fact is a mathematical function of the p-value

Power isn’t everything

  • It is not wrong to do an underpowered study, if it is the only way we can learn about a system
  • It may be prohibitively expensive to add more experimental units
  • It is especially important to be careful and precise with measurements and to ensure your small sample is representative
  • Draw conclusions with caution, and be explicit about the limitations of your inference

What software should I use?

R packages:

SAS procedures:

  • proc power
  • proc glmpower
  • proc glimmix

What did we learn today about power analysis?

  • Power analysis tells you how much you can trust the answers you get to the questions you care about
  • Determining effect size(s) is the most difficult but most important part of power analysis
  • Power analysis often requires you to simulate data based on assumptions
  • The more effort you put into a power analysis, the better informed your assumptions are by real data, the better your power analysis

Final recommendations

Image (c) Hasbro
  • Do low-powered studies at your own risk!
  • Design your study around the weakest effect
  • Think about justifying your design decisions to reviewers, readers, and yourself
  • Ask yourself “How confident will I be in the results of this study?”
  • Please don’t think of power analysis as a box to check, but as a critical part of the scientific method

No power = no new knowledge = no science