01 · The Question
Is “My Study Has a Larger Sample” Enough of a Research Contribution?
You find an existing study that asks almost exactly the question you want to investigate. It included 150 participants. You can recruit 1,000. Is that enough to justify doing the study again?
Possibly, but the difference in sample size is not itself the justification. A larger sample matters because of what those additional observations allow the new study to estimate, detect, distinguish, or represent more adequately.
If the original evidence is too imprecise, inadequately powered for an important effect, or unable to examine meaningful variation, a larger study may substantially improve the evidence. If the main problem is bias, poor measurement, confounding, or a weak design, simply adding participants may leave the important limitation almost untouched.
03 · What You Need to Know
Ask What the Larger Sample Will Let You Learn
Larger Samples Usually Improve Precision
One of the clearest benefits of increasing sample size is greater statistical precision. Estimates calculated from samples contain sampling uncertainty because a different sample drawn from the same population would produce somewhat different results.
Other things being equal, larger samples tend to produce more precise effect estimates and narrower confidence intervals. The Cochrane Handbook notes that the width of a confidence interval depends substantially on sample size, although variability in the outcome and, for some measures, the number or frequency of events also matter.
This gives you a stronger justification than simply saying, “The previous study had a small sample.” Ask instead whether its estimate was too imprecise for the question being asked.
Larger sample
Describes a difference in the amount of data collected.
Greater information
Describes what those additional data allow researchers to estimate or conclude more adequately.
The second is the scientific rationale. The first is merely a design characteristic.
A Larger Sample May Help When the Existing Study Was Underpowered
Statistical power is the probability that a study will detect an effect of a specified size under the assumptions used in the power calculation. Small studies may have insufficient power to detect effects that matter, particularly when those effects are modest or outcomes are highly variable.
A larger replication can therefore be useful when an earlier study was not adequately capable of detecting the effect that the research question requires it to distinguish.
This issue is particularly important when interpreting a previous nonsignificant finding. A nonsignificant result from a small study does not establish that there is no meaningful effect. The study may simply have produced an estimate with substantial uncertainty.
Research on replication sample sizes illustrates the point. Van Zwet and colleagues showed that replication studies may require considerably larger samples than the original studies to achieve high predictive power, particularly when the original evidence is only modestly statistically convincing. The precise multiplier depends on the assumptions and original result, so there is no general rule that a replication should simply double or triple the original sample.
A Larger Sample Can Distinguish Between Effect Sizes That Matter Differently
Power is not the only reason to collect more observations. Often, the more useful objective is estimation.
Imagine that an earlier study estimates an improvement of 5 points, but its confidence interval ranges from almost no improvement to an effect large enough to change practice. Knowing only the point estimate is not enough. The uncertainty encompasses substantively different interpretations.
A larger study may narrow that interval. If the new evidence can distinguish a trivial effect from one that is educationally, clinically, economically, or practically important, increased precision becomes a meaningful contribution.
This is closely related to improving certainty about an existing answer. The contribution may not be discovering a different effect. It may be learning the magnitude of the existing effect well enough to use the evidence responsibly.
A Larger Sample Can Support Analyses That Smaller Studies Could Not Reliably Address
Sometimes the primary estimate is not the unresolved issue. Researchers may need to know whether an effect varies across substantively important groups or conditions.
A study that is adequate for estimating an overall association may contain too few observations in relevant subgroups to estimate heterogeneity reliably. A larger sample could provide enough information to investigate a prespecified and theoretically justified interaction or subgroup question.
This does not mean that a large dataset licenses unlimited subgroup searching. As the number of exploratory comparisons increases, so does the opportunity to find unstable patterns by chance. The contribution is strongest when the additional sample enables analyses that were motivated before looking at the results and that answer questions the existing evidence could not address adequately.
Sample Size and Representativeness Are Different Problems
A very large sample can still represent the target population poorly.
Suppose an earlier study recruited 300 participants using a reasonably appropriate sampling strategy. A new researcher obtains 20,000 voluntary responses from users of one online platform. The second study is dramatically larger, but that does not automatically make its estimates more representative of the population of interest.
Sampling methodology matters. Probability sampling gives population members known selection probabilities and, when implemented appropriately, supports forms of population inference that a convenience sample may not. Nonprobability samples can still be useful for many research purposes, but increasing their size does not automatically remove selection bias.
Watch Out
A huge convenience sample is not automatically superior to a smaller, appropriately sampled dataset. Increasing N reduces sampling variability under the relevant assumptions; it does not guarantee that the people in the sample adequately represent the population you want to describe.
A Larger Sample Does Not Fix Bias
This distinction is fundamental. Sampling error and systematic bias are different sources of error.
Increasing sample size generally reduces sampling error. It does not necessarily reduce bias. Guidance on trial design makes this distinction explicitly: trial size can reduce sampling error, whereas bias generally must be addressed through design and conduct.
Imagine that a questionnaire systematically overestimates the construct it is supposed to measure. Administering that questionnaire to 10,000 people may estimate the biased quantity very precisely. You now have more precision around the wrong measurement.
The same logic applies to systematic selection problems, uncontrolled confounding, differential attrition, inappropriate comparison groups, and other design weaknesses. More observations can make estimates numerically stable without making the underlying inference more credible.
A Larger Sample Does Not Automatically Fix Confounding
This deserves particular attention because researchers sometimes assume that a large dataset can compensate for a weak observational design.
It cannot do so merely by being large. Confounding occurs when differences related to both an exposure and an outcome create an alternative explanation for an observed association. Larger samples may permit more sophisticated adjustment when appropriate confounders have been measured, but they do not magically measure unobserved confounders or guarantee that an analytical adjustment is valid.
If confounding is the principal reason the existing literature cannot support the desired inference, a better-controlled study may be more important than a merely larger one.
A Larger Sample Does Not Repair Poor Measurement
Measurement error creates another situation in which “more” can be mistaken for “better.”
If an existing study uses an instrument that poorly captures the construct of interest, collecting a much larger sample with the same problematic measure may reproduce the same limitation more precisely.
The appropriate improvement may instead involve better measurement, perhaps alongside a larger sample when both problems matter.
Research design improvements should therefore be matched to the weakness in the evidence. Sample size is not a universal methodological repair kit.
Very Large Samples Can Make Trivial Effects Statistically Significant
As sample size increases, statistical power increases. Consequently, sufficiently large studies can detect effects that are very small.
This is useful when small effects genuinely matter. It becomes problematic when statistical significance is interpreted as evidence that an effect is important.
A large study may provide compelling evidence that an effect differs from zero while simultaneously showing that the effect is too small to matter for the research problem. Researchers should therefore examine effect estimates, confidence intervals, and substantive importance rather than treating a smaller p-value as the primary reward for increasing sample size.
The Existing Evidence Determines Whether More Precision Is Worthwhile
There is no inherent scientific virtue in making every estimate increasingly precise.
If several rigorous studies already estimate an effect narrowly enough to support the relevant conclusion, another much larger study may produce only a marginal reduction in uncertainty. At that point, the same resources might generate more useful knowledge by addressing a different unresolved issue.
Before proposing a larger replication, determine whether the existing evidence is already sufficiently informative. If it is not, identify exactly how much uncertainty remains and whether the proposed sample is capable of reducing it meaningfully.
The Required Sample Should Follow the Research Objective
“Bigger than the previous study” is not a sample-size calculation.
The appropriate sample depends on what the study is designed to accomplish. A study powered to detect a prespecified effect requires assumptions about the effect, variability, significance criterion, desired power, design, and other relevant features. A study designed primarily for estimation may instead be planned around a desired level of precision, such as an acceptable confidence-interval width.
Replication presents additional complications because the effect reported in an original study may overestimate the effect likely to be observed again. Simply calculating the replication sample from the original point estimate can therefore be optimistic.
The methodological details vary across designs and disciplines. The underlying principle does not: determine the information you need first, then determine the sample required to obtain it.