03 · What You Need to Know
Why More Observations Cannot Automatically Overcome Systematic Bias
Sample Size and Sample Quality Answer Different Questions
Sample size tells you how many observations contribute information to an analysis. It does not tell you whether those observations entered the study through a process appropriate to the population and inference you care about.
A sample of 50,000 volunteers and a probability sample of 2,000 people are not distinguished merely by the fact that one contains 25 times as many observations. They are generated through different selection mechanisms.
If the research objective is population estimation, that mechanism matters because the observed participants need to provide a defensible basis for learning about people who were not observed.
This is why asking whether a bigger sample automatically makes a study better cannot be answered by participant count alone.
Random Error Usually Shrinks as Sample Size Grows
Under common statistical conditions, increasing the sample size reduces the standard error of an estimate. For a simple mean based on independent observations, the familiar relationship is:
The square-root relationship explains why larger samples can produce narrower confidence intervals. It also reveals something important: this calculation concerns sampling variability under a statistical model. It does not contain a term that magically corrects undercoverage, self-selection, nonresponse, or poor measurement.
Bias Does Not Have to Shrink With the Standard Error
Suppose the true population prevalence of a behavior is 40%, but the recruitment process systematically overrepresents people who engage in that behavior, causing the sampling process to produce estimates centered near 55%.
Increasing the sample may make the estimate settle more tightly around approximately 55% rather than moving it toward 40%.
This is the uncomfortable possibility behind large-sample bias: the estimate becomes more stable without becoming more accurate for the target population.
High precision
Repeated estimates under the sampling process would be tightly concentrated.
Low bias
The estimation process is centered appropriately on the target quantity under the relevant assumptions.
A study can have one without the other.
The Big Data Paradox Makes This Problem Particularly Visible
Statistician Xiao-Li Meng used the term Big Data Paradox to describe a striking consequence of systematic data defects: as datasets become larger, seemingly negligible correlations between inclusion in the dataset and the quantity being measured can dominate random sampling error.
The implication is counterintuitive. Very large datasets are not automatically protected from selection problems merely because they include a substantial number of observations. When data quantity grows much faster than data representativeness improves, conventional measures of sampling precision can create an increasingly misleading sense of certainty.
The lesson is not that big data are inherently unreliable. It is that data quantity and data quality are mathematically different dimensions of evidence.
Undercoverage Does Not Disappear When You Sample More From the Covered Population
Suppose your target population contains 100,000 instructors, but your sampling frame systematically excludes 20,000 part-time instructors.
You could sample 500, 5,000, or every single instructor appearing on the frame. None of those choices directly provides observations from the 20,000 eligible instructors who are absent.
If part-time instructors differ from covered instructors on the outcome, the population estimate may remain biased.
The appropriate response is to examine the mismatch between the sampling frame and population, not merely to increase the number selected from the incomplete frame.
Self-Selection Can Scale Up Along With the Sample
Open online surveys can attract enormous numbers of respondents. That is useful for many research purposes, but the mechanism by which people volunteer still matters.
Suppose a questionnaire about academic use of generative AI is promoted heavily through AI enthusiast communities. People with strong interest, experience, or opinions may be more likely to encounter the invitation and participate.
If 1,000 such volunteers are selective, recruiting 100,000 through essentially the same mechanism does not automatically make the sample representative. The selection mechanism has simply operated at a larger scale.
This is one reason convenience sampling should be evaluated according to how accessibility relates to the research question rather than according to the final respondent count.
Nonresponse Can Create a Large but Selective Achieved Sample
Large initial samples do not guarantee that the final responding sample preserves the intended design.
Imagine inviting 200,000 people and obtaining 40,000 responses. Forty thousand is still a very large sample. Yet if participation is systematically related to the outcome, the responding sample may produce biased estimates despite its size.
This is why nonresponse bias depends on who responds and how respondents differ from nonrespondents, not simply on whether the final dataset looks numerically impressive.
A Huge Sample Can Have a Tiny Standard Error and a Large Total Error
Researchers sometimes report very narrow confidence intervals from enormous non-probability datasets and interpret those intervals as evidence that the population estimate must be highly accurate.
Conventional confidence intervals typically quantify uncertainty under a specified statistical model and sampling assumptions. They do not automatically incorporate unknown systematic errors arising from selection, coverage, measurement, or model misspecification.
A narrow interval can therefore coexist with substantial uncertainty about whether the estimand corresponds to the intended target population.
Watch Out
Do not interpret a narrow confidence interval as proof that selection bias is small. The interval may quantify random uncertainty while leaving important systematic uncertainty outside the calculation.
Large Samples Can Make Trivial Effects Look Impressive
Large samples also change hypothesis testing. As standard errors shrink, very small departures from a null value can become statistically significant.
A difference may therefore produce a very small p-value while remaining substantively negligible.
This is a separate issue from sampling bias, but the two can combine awkwardly. A huge selectively sampled dataset may produce an extremely precise and highly statistically significant association that is both systematically distorted and practically unimportant.
Effect magnitude, uncertainty, research design, and substantive relevance still require interpretation.
Does a Small Probability Sample Always Beat a Large Non-Probability Sample?
No. That would merely replace one simplistic rule with another.
A well-designed probability sample offers a formal design-based foundation for population inference, but it may itself suffer from coverage problems, nonresponse, measurement error, or inadequate sample size. A large non-probability dataset may contain valuable information and, with appropriate auxiliary data, assumptions, and statistical methods, may sometimes support useful population inference.
The point is not that smaller is better. The point is that size cannot substitute for understanding how the data came to represent the target population.
Can Weighting Rescue a Very Large Biased Sample?
Weighting, calibration, matching, propensity modeling, and data-integration methods can sometimes reduce differences between a sample and a target population.
Their success depends on the quality of external population information, measurement comparability, model specification, and whether the variables needed to address selection have actually been observed.
If selection depends strongly on an unmeasured characteristic related to the outcome, adjusting perfectly for age, sex, region, and education may still leave consequential bias.
Large sample size does not remove these assumptions. Indeed, when random error is already tiny, uncertainty about those assumptions may become more important than collecting another few thousand observations.
What Should You Examine Instead of Sample Size Alone?
Start with the target population and trace the path by which observations entered the dataset.
Who was covered? Who could be selected? Who encountered recruitment? Who volunteered or responded? Who remained in the analysis? What adjustments were applied? Which differences between the observed sample and target population are measurable, and which remain unknown?
These questions connect sample size with the broader issue of sampling bias. A million observations do not make those questions obsolete. They make answering them more consequential.