03 · What You Need to Know
What Actually Determines Quantitative Sample Size?
Start With the Primary Research Objective
Before calculating sample size, identify the primary quantity, comparison, association, or effect the study must estimate or test.
Are you estimating the proportion of university students who use generative AI? Comparing examination scores between two instructional conditions? Testing an association between AI literacy and academic self-efficacy? Estimating a regression coefficient? Evaluating whether an intervention produces a clinically or educationally meaningful difference?
These are different statistical problems. They do not necessarily use the same sample-size formula or require the same number of observations.
Sample size should therefore follow the research question and analysis rather than precede them. Choosing a participant number first and searching afterward for a statistical justification rather misses the point of prospective sample-size planning.
What Is Statistical Power?
In conventional null-hypothesis significance testing, statistical power is the probability that a statistical test rejects the null hypothesis when a specified alternative is true.
If a study has 80% power for a particular effect under specified assumptions, that means that, over hypothetical repetitions of the study under those assumptions, the testing procedure would reject the null hypothesis about 80% of the time when that specified alternative is true.
Power is therefore not the probability that your hypothesis is correct. It is also not a general score describing the quality of a study.
Statistical power
The probability of rejecting the null hypothesis under a specified alternative, given the design and assumptions.
Statistical significance
A result of applying a statistical decision rule to the observed data, commonly by comparing a p-value with a prespecified significance threshold.
Four Quantities Are Closely Connected in a Basic Power Analysis
For many common hypothesis tests, prospective power analysis involves a relationship among the effect size of interest, significance level, desired power, and sample size. Once appropriate values for the other quantities are specified, sample size can be solved under the assumed statistical model.
Effect Size Is Often the Hardest Assumption
A power analysis needs an effect to power the study for. That assumption should not be chosen merely because a conventional label calls an effect “small,” “medium,” or “large.”
The more useful question is: What is the smallest effect that would be substantively important enough for this study to detect?
Prior studies, meta-analyses, pilot data, theory, domain expertise, or a clearly argued smallest effect size of interest may help inform this choice. Each source has limitations. Pilot studies, for example, can produce unstable effect estimates when they are small, while published effects may be affected by differences in population, design, measurement, or publication processes.
A sample-size calculation can be mathematically exact while resting on an implausible effect-size assumption. The assumptions deserve at least as much attention as the software output.
Smaller Effects Usually Require Larger Samples
Imagine two educational interventions. One study is designed only to detect a very large difference in learning outcomes. Another needs adequate power to detect a much smaller but educationally meaningful difference.
All else being equal, the second study generally requires a larger sample because smaller signals are more difficult to distinguish from sampling variability.
This is one reason simply copying the sample size from a previous study is weak justification. The earlier study may have targeted a different effect, used a different outcome, or employed a different design.
Higher Desired Power Usually Requires More Participants
If all other assumptions remain unchanged, requiring a higher probability of detecting the specified effect generally increases the required sample size.
Values such as 80% or 90% power are commonly used in research planning, but they are conventions rather than laws of nature. The appropriate choice should reflect the consequences of failing to detect an effect, disciplinary expectations, resources, and the study's purpose.
Reporting “power =.80” without explaining the effect size, significance level, statistical test, and other assumptions leaves the calculation difficult to evaluate.
The Significance Level Also Affects Sample Size
The significance level, commonly denoted by α, sets the Type I error probability under the statistical testing framework. A conventional value such as.05 is frequently used, but it should not be inserted into a calculation without understanding its role.
Holding other factors constant, using a more stringent significance threshold generally requires more information to achieve the same power for a specified effect.
Multiple primary comparisons can further complicate planning when the design controls an overall error rate across tests. Sample-size calculations should correspond to the actual testing strategy rather than an imaginary single test that will never appear in the analysis.
Not Every Study Should Be Planned Around Statistical Power
Suppose your primary objective is to estimate the percentage of teachers who use generative AI, with a confidence interval narrow enough to be useful. Your central concern is then precision of estimation, not necessarily power for rejecting a null hypothesis.
For a proportion under a simple random sampling approximation, planning may be expressed through the desired margin of error:
The familiar figure of roughly 384 therefore is not a universal sample size for surveys. It is the result of a particular calculation under particular assumptions. Change the desired precision, expected proportion, confidence level, or sampling design and the requirement changes.
Population Size Sometimes Matters Less Than Researchers Expect
A common intuition is that a population of one million must always require a dramatically larger sample than a population of ten thousand.
For estimating proportions under simple random sampling, once the population is large relative to the sample, required sample size is driven more strongly by desired precision, confidence level, and variability than by population size itself.
When the planned sample is a substantial fraction of a finite population, however, a finite population correction may become relevant. The details depend on the design and inferential framework.
This is why rules such as “sample 10% of the population” are not general substitutes for sample-size calculation.
Complex Sampling Designs Can Change the Requirement
The simplest calculations often assume independent observations obtained through simple random sampling. Real studies may use stratification, clustering, unequal probabilities, repeated observations, or multiple stages.
Clustered designs are especially important. Participants within the same school, classroom, hospital, or community may resemble one another, so 500 clustered observations may contain less independent information than 500 independently sampled observations.
For equal-sized clusters under a simple approximation, the inflation in variance is often described using a design effect:
If you are using stratified, cluster, systematic, or multistage sampling, the sample-size plan should reflect the actual design.
Plan for Nonresponse, Attrition, and Unusable Data
The number required for analysis is not necessarily the number you should recruit.
Survey invitations may go unanswered. Participants in longitudinal studies may withdraw. Some observations may fail eligibility checks or lack information needed for the primary analysis.
If you require a final analyzable sample of 400 and reasonably expect 80% of recruited participants to provide usable data, simply recruiting 400 would be inadequate.
The anticipated loss rate should be evidence-informed where possible rather than selected merely to inflate the target “just in case.”
Subgroup Analyses May Determine the Real Sample-Size Requirement
A study may have enough participants overall but too few within important subgroups.
Suppose you need to compare four academic disciplines but one discipline represents only 5% of the population. A sample that is adequate for an overall estimate may contain too few participants from that subgroup for the planned comparison.
This can affect both sample size and sampling design. Stratified sampling or oversampling may be appropriate when subgroup estimates are genuine research objectives.
More Participants Do Not Automatically Make a Better Study
Increasing sample size can improve precision and power under appropriate conditions, but it does not correct a poorly defined population, biased recruitment process, invalid measurement, confounding, inappropriate analysis, or flawed study design.
A very large sample can also make tiny effects statistically detectable even when those effects have little substantive importance.
This is why the question of whether a bigger sample automatically makes research better must be separated from the narrower statistical question of how sample size affects uncertainty and power.