01 · The Question
Is a Small Sample Size Enough to Dismiss a Study?
You are reading a paper and notice that the study included only 30 participants. Or perhaps 18 interviews, 12 schools, six laboratories, or four case sites. The number looks small, and the temptation is immediate: the study must be weak.
Sometimes that concern is justified. An insufficient sample can make estimates imprecise, leave a quantitative study with too little statistical power, or provide too little evidence for the conclusions being drawn. But the number of observations alone does not tell you whether any of those problems actually occurred.
Sample adequacy depends on what the researchers were trying to learn, how the study was designed, what kind of data were collected, how the data were analyzed, and how ambitious the conclusions are. The more useful question is therefore not simply “Is this sample small?” but “Is this sample adequate for what this study is trying to claim?”
03 · What You Need to Know
What Actually Determines Whether a Sample Is Too Small?
There Is No Universal Number That Separates “Small” From “Large”
There is no defensible rule that says a study becomes strong at 100 participants, 200 respondents, or any other universal threshold. A sample that is entirely reasonable for one research design may be inadequate for another.
For example, a tightly controlled within-participant experiment may require fewer participants than a study attempting to estimate a small association precisely across a heterogeneous population. A rare-disease study may face recruitment constraints that do not exist in a large online survey. A qualitative interview study may deliberately use a relatively small purposive sample because its purpose is intensive analysis rather than statistical estimation of population parameters.
This is why evaluating whether a sample was appropriate requires more than counting participants.
For Quantitative Studies, Ask What the Sample Allows the Researchers to Detect or Estimate
In many quantitative studies, sample size directly affects statistical power and precision. Statistical power is the probability that a statistical procedure will detect an effect of a specified size when that effect is present under the assumptions of the analysis. Other things being equal, larger samples generally provide greater power.
A study with low power may fail to detect an effect that matters. This is especially important when authors interpret a non-significant result as evidence that there is no effect. An underpowered study may simply have been unable to distinguish a meaningful effect from sampling variation. Low power has also been associated with unstable effect estimates and exaggerated estimates among statistically significant findings.
But power is not determined by sample size alone. It also depends on factors such as the effect size being investigated, the statistical model, variability in the data, significance criterion, study design, and measurement reliability. A particular sample size therefore cannot be labeled “underpowered” merely because the number looks modest.
Precision May Matter More Than Whether a Result Is Statistically Significant
Sample size also affects the precision of estimates. With less information, confidence intervals or other uncertainty intervals will generally be wider, all else being equal. That means the study may be compatible with a relatively broad range of plausible effect sizes.
Suppose a small trial estimates that an intervention improves an outcome by 5 points, but its confidence interval is compatible with anything from a negligible improvement to a substantial one. The problem is not simply that the study had “few participants.” The more informative criticism is that the estimate is too imprecise to support a narrow conclusion about the intervention's likely effect.
This distinction matters because a statistically significant estimate can still be imprecise, while a non-significant estimate can sometimes provide useful information if its uncertainty is sufficiently narrow for the research question.
Sample Size Should Be Judged Against the Effect That Actually Matters
Power calculations depend partly on the effect size researchers want to be able to detect. That assumption deserves scrutiny.
A study can appear adequately powered if researchers assume an unrealistically large effect. If the effects that would matter scientifically or practically are smaller, the same sample may be inadequate. Contemporary statistical guidance therefore increasingly emphasizes explicit justification of sample-size assumptions and consideration of the smallest effect that would be substantively meaningful rather than mechanically choosing a conventional “small,” “medium,” or “large” effect.
When critically reading a paper, ask what effect size the sample-size calculation assumed, where that assumption came from, and whether an effect smaller than that would still matter.
A Sample-Size Calculation Is Useful, but It Is Not a Certificate of Quality
An a priori power analysis can provide a principled rationale for sample size in many confirmatory quantitative designs. Yet seeing the words “power analysis” in the Methods section should not end your evaluation.
Look at the assumptions. What analysis was used for the calculation? What effect size was assumed? What power level and significance threshold were chosen? Was attrition anticipated? If several important analyses were planned, was the calculation appropriate for the analysis that required the most information?
A technically correct calculation based on unrealistic assumptions can still justify an inadequate sample. Conversely, some designs use other defensible approaches to sample-size planning, including precision-based planning, simulation, sequential designs, or constraints imposed by rare populations and specialized data collection.
The Number Recruited May Not Be the Number That Actually Matters
A paper may report that 200 participants were recruited, but the effective information available for a particular analysis can be much smaller. Attrition, missing observations, exclusions, clustering, repeated measurements, unequal group sizes, or subgroup analyses can all change how much information an analysis contains.
Imagine a study with 120 participants divided among four experimental conditions. The headline sample is 120, but the relevant comparisons may involve roughly 30 participants per condition. If the paper then conducts an analysis within one subgroup, the effective sample for that claim may shrink further.
For clustered research, 500 students from only five schools also do not provide the same independent information as 500 independently sampled students. Participants within the same school may resemble one another, so the number of clusters can become critically important.
A Small Sample Cannot Be Rescued by Poor Design
Sometimes discussions of sample size distract from more consequential problems. A larger sample cannot repair invalid measurement, severe selection bias, uncontrolled confounding, inappropriate analysis, or a design incapable of answering the research question.
The reverse is also important. A modest sample does not erase the strengths of careful measurement, rigorous experimental control, transparent procedures, or a design well matched to the question. You still need to evaluate whether the study design can answer the research question, whether the measures are adequate, and whether the analysis fits the design.
Representativeness and Sample Size Are Different Problems
A large sample is not necessarily representative, and a smaller sample is not necessarily badly selected. Sample size concerns how much information is available; sampling strategy concerns, among other things, who or what contributed that information and what population the findings may reasonably describe.
Sample size
How many relevant observations, participants, cases, clusters, or other units contribute information to the study and its analyses.
Sample representativeness
How well the sampled units support inference to the population the researchers intend to describe or generalize to, given the sampling process and other sources of selection.
Increasing the number of observations does not automatically remove selection bias. Ten thousand poorly selected respondents can provide a very precise estimate of the wrong population quantity.
Qualitative Research Requires a Different Logic
Applying conventional quantitative power rules to qualitative research is usually inappropriate. A qualitative study may intentionally examine a small number of participants or cases because it seeks depth, contextual understanding, theoretical development, or detailed interpretation rather than statistical estimation of population effects.
Qualitative sample adequacy should instead be evaluated in relation to the methodology, research aim, sampling strategy, richness and relevance of the data, analytic approach, and the claims being made. One influential formulation describes this in terms of information power: the more relevant information a sample provides for the study aim, the fewer participants may be required. Factors include the specificity of the sample, use of theory, quality of dialogue, and analysis strategy.
That does not mean sample size is irrelevant in qualitative research. A sample can still be too limited for the analytic claims being made. It means that qualitative research should be evaluated using standards appropriate to qualitative inquiry rather than by importing quantitative thresholds.
Small Samples Become More Concerning When the Claims Become More Ambitious
The adequacy of a sample is inseparable from the scope of the conclusion. A small exploratory study may provide useful preliminary evidence, demonstrate feasibility, document an unusual phenomenon, or generate hypotheses without claiming to establish a stable population effect.
The same sample becomes more problematic if the authors make highly precise estimates, claim that an effect is absent, generalize broadly to heterogeneous populations, conduct many subgroup analyses, or present an unstable result as definitive.
When reading the Discussion and Conclusion, therefore, examine whether the authors have calibrated their language to the amount and quality of evidence available. This is part of recognizing when researchers are overstating what their results support.
07 · A Quick Checklist
What to Check Before Calling a Sample Too Small
Before judging the study, check:
What research question and primary claim is the sample expected to support?
Did the authors explain how the sample size was determined or otherwise justify why it was adequate?
For quantitative inference, what effect size, precision target, statistical power, significance criterion, or other assumptions informed sample planning?
How wide are the confidence intervals or other uncertainty estimates around the main findings?
Did attrition, missing data, exclusions, unequal groups, clustering, or subgroup analyses reduce the information available for important analyses?
Does the sampling strategy support the population or setting to which the authors generalize?
Are non-significant findings being interpreted cautiously rather than treated automatically as evidence that no meaningful effect exists?
For qualitative research, is sample adequacy justified using criteria appropriate to the methodology and analytic purpose?
Are the authors' conclusions appropriately limited to what the available sample and design can support?