03 · What You Need to Know
Sample Size Solves Some Problems, Not Every Research Problem
The attraction of a large sample is statistically sensible. Estimates based on more information are generally less affected by sampling variability, all else being equal. Confidence intervals often become narrower, and researchers may be able to detect smaller effects.
The phrase “all else being equal” does considerable work, however. A large study is not simply a small study made better. Its usefulness depends on where the observations came from, what was measured, how the study was designed, what assumptions were made, and what conclusion is being drawn.
Large samples primarily improve precision
Imagine estimating the average height of students at a university. If participants are sampled appropriately and measured accurately, an estimate based on 5,000 students will ordinarily be more precise than one based on 50. Repeated random samples of 5,000 students would tend to fluctuate less.
This reduction in sampling variability is a genuine advantage. It is one reason a large, well-designed study can sometimes be especially informative.
But precision describes how much an estimate would tend to vary under the study's sampling and statistical framework. It does not establish that the estimate is centered on the quantity researchers actually intend to estimate.
Precision
How tightly an estimate is determined given the information and statistical model used.
Validity
Whether the study design, measurements, analysis, and interpretation support an estimate or conclusion that adequately corresponds to the intended research question.
A result can therefore be precise but misleading. A narrow confidence interval around a biased estimate does not make the bias disappear.
A huge unrepresentative sample can remain unrepresentative
One of the clearest examples involves sampling. Suppose researchers want to estimate a characteristic of an entire population but collect data disproportionately from people who differ systematically from those who are missing.
Increasing the number of respondents within that selected group can make the estimate increasingly stable while leaving the selection problem intact.
Kaplan and colleagues, writing about big data and large sample sizes, emphasize that representativeness can matter more than sheer sample size for many inferences. They use the well-known 1936 Literary Digest presidential poll as an illustration: millions of responses did not rescue a sampling process that produced an unrepresentative picture of the electorate.
The historical lesson is not that large samples are undesirable. It is that a huge sample from the wrong sampling process can be inferior to a smaller sample designed to represent the target population more appropriately.
Selection bias does not disappear when more selected participants are added
Sampling problems extend beyond surveys. Participation in health databases, educational platforms, online experiments, voluntary registries, and observational studies may depend on characteristics related to the outcome being studied.
If selection into the dataset systematically differs across relevant groups, increasing the number of selected observations does not necessarily recreate the observations that are missing.
A million volunteers are still volunteers. A million users of one platform do not automatically represent people who never use the platform. A national database may be enormous while systematically excluding particular populations.
Measurement error can become very precisely reproduced
A study cannot recover information that its measurement process fails to capture merely by adding participants.
Suppose researchers estimate physical activity from a systematically biased self-report measure. A larger sample can improve the precision of statistics calculated from those reported values. It does not automatically make the reports accurate measurements of actual activity.
Methodological research on measurement error makes this distinction explicit: increasing sample size can improve precision around the expected estimate under the measurement-error process without necessarily bringing that estimate closer to the true value.
The same principle applies to poorly operationalized constructs, misclassified exposures, unreliable administrative codes, incomplete records, and instruments that systematically miss part of the phenomenon being investigated.
Large observational studies can still be confounded
Suppose a study of one million people finds that people who engage in behavior A have better outcome B. The sample is enormous, and the association is estimated with extraordinary precision.
If people who engage in A systematically differ in other relevant ways from those who do not, the association may still reflect confounding. Statistical adjustment can help when relevant confounders are appropriately measured and modeled, but sample size itself cannot adjust for an important variable that was never adequately captured.
This is a fundamental reason large observational studies should not be interpreted as randomized experiments merely because their datasets are enormous.
Watch Out
A very narrow confidence interval does not account automatically for every source of uncertainty. It may quantify sampling uncertainty under the statistical model while leaving important uncertainty from bias, confounding, measurement, model specification, or selection outside the interval.
Large samples can make tiny differences statistically significant
Statistical significance is influenced by sample size. With enough observations, a study may obtain a small p-value for an association or difference that is extremely modest in magnitude.
This is not a defect in significance testing so much as a common error in interpretation. A small p-value does not tell you that an effect is large, important, useful, or causal.
Current guidance on epidemiological evidence emphasizes this distinction: a very large study may detect a statistically significant association of modest magnitude, whereas a smaller study may estimate a much larger association but remain statistically uncertain.
Researchers should therefore inspect effect estimates, confidence intervals, practical relevance, and study validity rather than translating “statistically significant in a huge study” into “important and definitely true.”
Big datasets can contain many opportunities to find patterns
Large datasets often include many variables, outcomes, subgroups, time periods, and possible model specifications. This richness creates scientific opportunities, but it also creates analytical flexibility.
If researchers explore enough relationships without appropriate safeguards, some apparently compelling patterns may emerge through chance, selective analysis, or selective reporting. The problem is not solved simply because every analysis contains millions of observations.
Pre-specification when appropriate, transparent reporting, multiplicity adjustments where warranted, robustness analyses, and independent confirmation can help distinguish findings that survive scrutiny from patterns that emerged from a large analytical search space.
Missing information can matter more than the information you have
A dataset can contain millions of complete records for some variables while systematically lacking the variable most important for interpreting the association.
For example, researchers may have detailed records of educational platform use and examination scores but inadequate information about prior achievement, motivation, instructor differences, or why students chose to use the platform. The number of rows does not compensate automatically for the absence of the variable needed to address the competing explanation.
Large-scale data are especially powerful when the variables needed for the research question are measured appropriately. Volume cannot substitute for relevance.
A large study may answer a narrow question extremely well
Sometimes the problem is not that the estimate is inaccurate, but that researchers generalize beyond what the study actually establishes.
A study involving 500,000 adults in one healthcare system may precisely describe patterns within that system. Whether the same estimates apply to children, people in other countries, uninsured populations, or substantially different healthcare settings is a separate question.
Sample size within one population does not automatically create diversity beyond that population.
More precision can sometimes increase false confidence
This is the counterintuitive part. When a biased study is small, wide uncertainty may make readers cautious. When the same systematic problem occurs in a massive dataset, the estimate can appear exceptionally stable.
Researchers may then mistake statistical precision for epistemic certainty.
Kaplan and colleagues caution that large samples can magnify inferential problems associated with sampling and study-design errors. The concern is not that additional observations literally create every bias, but that large datasets can produce highly persuasive-looking statistics while the systematic problem remains.
This is closely connected to why more evidence does not automatically mean better evidence. Quantity is valuable when the information being accumulated is capable of answering the intended question.
| Problem |
Does a larger sample help? |
What else is needed? |
| Sampling variability |
Usually yes |
Sufficient relevant observations and an appropriate statistical analysis |
| Low precision |
Often yes |
Enough informative observations for the estimate of interest |
| Unrepresentative sampling |
Not automatically |
A sampling strategy or adjustment capable of addressing the selection problem |
| Systematic measurement error |
Not automatically |
Better measurement, validation, or appropriate methods addressing measurement error |
| Uncontrolled confounding |
Not by size alone |
A design or analysis capable of addressing the competing explanation |
| Tiny but statistically detectable effect |
May make detection easier |
Interpretation of effect magnitude and practical importance |
| Narrow population or setting |
Not automatically |
Evidence about whether the conclusion applies elsewhere |
| Missing important variables |
Usually not |
Relevant measurements or a design that addresses the missing information |
A smaller study can sometimes be more informative
A smaller, purpose-designed study may collect a more representative sample, measure variables more carefully, implement stronger controls, or obtain information unavailable in a massive secondary dataset.
That does not make smaller studies inherently better. Small studies face their own problems, especially limited precision and difficulty detecting modest effects. The comparison simply illustrates that study size is one methodological property among several.
Indeed, a small study can sometimes be scientifically important precisely because it contributes information that a larger study cannot easily obtain.