Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can a Sample Be Large but Still Seriously Biased?

A large sample can reduce random uncertainty while leaving systematic bias almost untouched. Learn why thousands or even millions of observations can produce extremely precise answers to the wrong population question.

94
Can a Large Sample Still Be Biased? Guide 94 of 217
01 · The Question

If Your Sample Is Huge, Can It Still Give You the Wrong Answer?

Imagine surveying 100,000 university students about generative AI use. Compared with a study of 500 students, the larger dataset looks formidable. Its percentages may appear stable, its confidence intervals may be narrow, and even small differences between groups may be statistically detectable.

But suppose nearly all 100,000 participants came from online communities devoted to artificial intelligence.

Now the sample size looks rather less reassuring.

This illustrates an important statistical distinction: a large sample can reduce random uncertainty without eliminating systematic bias. If the process producing the sample disproportionately includes some parts of the target population and excludes others, increasing the number of observations through that same process may give you a more precise estimate without giving you a more accurate estimate of the population quantity you actually care about.

02 · The Short Answer

Large Samples Can Be Extremely Precise and Still Seriously Wrong

In Brief

Yes. A sample can contain thousands or millions of observations and still produce seriously biased population estimates when coverage, selection, recruitment, nonresponse, measurement, or other systematic processes make the observed cases differ consequentially from the target population.

Increasing sample size primarily reduces random sampling variability under appropriate conditions. It does not automatically reduce systematic error, so a very large biased sample can produce narrow confidence intervals around an estimate that remains meaningfully displaced from the population value.

03 · What You Need to Know

Why More Observations Cannot Automatically Overcome Systematic Bias

Sample Size and Sample Quality Answer Different Questions

Sample size tells you how many observations contribute information to an analysis. It does not tell you whether those observations entered the study through a process appropriate to the population and inference you care about.

A sample of 50,000 volunteers and a probability sample of 2,000 people are not distinguished merely by the fact that one contains 25 times as many observations. They are generated through different selection mechanisms.

If the research objective is population estimation, that mechanism matters because the observed participants need to provide a defensible basis for learning about people who were not observed.

This is why asking whether a bigger sample automatically makes a study better cannot be answered by participant count alone.

Random Error Usually Shrinks as Sample Size Grows

Under common statistical conditions, increasing the sample size reduces the standard error of an estimate. For a simple mean based on independent observations, the familiar relationship is:

Sampling Precision
SE(x̄) = σ / √n
SE(x̄) is the standard error of the sample mean, σ is the population standard deviation under the model, and n is the sample size.
If σ = 20, a sample of 100 gives an idealized standard error of 20 / √100 = 2. Increasing the sample to 10,000 gives 20 / √10,000 = 0.2. Under the assumptions of this simple calculation, the larger sample estimates its sampling-process expectation much more precisely. The formula does not say that systematic selection or measurement bias has disappeared.

The square-root relationship explains why larger samples can produce narrower confidence intervals. It also reveals something important: this calculation concerns sampling variability under a statistical model. It does not contain a term that magically corrects undercoverage, self-selection, nonresponse, or poor measurement.

Bias Does Not Have to Shrink With the Standard Error

Suppose the true population prevalence of a behavior is 40%, but the recruitment process systematically overrepresents people who engage in that behavior, causing the sampling process to produce estimates centered near 55%.

Increasing the sample may make the estimate settle more tightly around approximately 55% rather than moving it toward 40%.

This is the uncomfortable possibility behind large-sample bias: the estimate becomes more stable without becoming more accurate for the target population.

High precision Repeated estimates under the sampling process would be tightly concentrated.
Low bias The estimation process is centered appropriately on the target quantity under the relevant assumptions.

A study can have one without the other.

The Big Data Paradox Makes This Problem Particularly Visible

Statistician Xiao-Li Meng used the term Big Data Paradox to describe a striking consequence of systematic data defects: as datasets become larger, seemingly negligible correlations between inclusion in the dataset and the quantity being measured can dominate random sampling error.

The implication is counterintuitive. Very large datasets are not automatically protected from selection problems merely because they include a substantial number of observations. When data quantity grows much faster than data representativeness improves, conventional measures of sampling precision can create an increasingly misleading sense of certainty.

The lesson is not that big data are inherently unreliable. It is that data quantity and data quality are mathematically different dimensions of evidence.

Undercoverage Does Not Disappear When You Sample More From the Covered Population

Suppose your target population contains 100,000 instructors, but your sampling frame systematically excludes 20,000 part-time instructors.

You could sample 500, 5,000, or every single instructor appearing on the frame. None of those choices directly provides observations from the 20,000 eligible instructors who are absent.

If part-time instructors differ from covered instructors on the outcome, the population estimate may remain biased.

The appropriate response is to examine the mismatch between the sampling frame and population, not merely to increase the number selected from the incomplete frame.

Self-Selection Can Scale Up Along With the Sample

Open online surveys can attract enormous numbers of respondents. That is useful for many research purposes, but the mechanism by which people volunteer still matters.

Suppose a questionnaire about academic use of generative AI is promoted heavily through AI enthusiast communities. People with strong interest, experience, or opinions may be more likely to encounter the invitation and participate.

If 1,000 such volunteers are selective, recruiting 100,000 through essentially the same mechanism does not automatically make the sample representative. The selection mechanism has simply operated at a larger scale.

This is one reason convenience sampling should be evaluated according to how accessibility relates to the research question rather than according to the final respondent count.

Nonresponse Can Create a Large but Selective Achieved Sample

Large initial samples do not guarantee that the final responding sample preserves the intended design.

Imagine inviting 200,000 people and obtaining 40,000 responses. Forty thousand is still a very large sample. Yet if participation is systematically related to the outcome, the responding sample may produce biased estimates despite its size.

This is why nonresponse bias depends on who responds and how respondents differ from nonrespondents, not simply on whether the final dataset looks numerically impressive.

A Huge Sample Can Have a Tiny Standard Error and a Large Total Error

Researchers sometimes report very narrow confidence intervals from enormous non-probability datasets and interpret those intervals as evidence that the population estimate must be highly accurate.

Conventional confidence intervals typically quantify uncertainty under a specified statistical model and sampling assumptions. They do not automatically incorporate unknown systematic errors arising from selection, coverage, measurement, or model misspecification.

A narrow interval can therefore coexist with substantial uncertainty about whether the estimand corresponds to the intended target population.

Watch Out

Do not interpret a narrow confidence interval as proof that selection bias is small. The interval may quantify random uncertainty while leaving important systematic uncertainty outside the calculation.

Large Samples Can Make Trivial Effects Look Impressive

Large samples also change hypothesis testing. As standard errors shrink, very small departures from a null value can become statistically significant.

A difference may therefore produce a very small p-value while remaining substantively negligible.

This is a separate issue from sampling bias, but the two can combine awkwardly. A huge selectively sampled dataset may produce an extremely precise and highly statistically significant association that is both systematically distorted and practically unimportant.

Effect magnitude, uncertainty, research design, and substantive relevance still require interpretation.

Does a Small Probability Sample Always Beat a Large Non-Probability Sample?

No. That would merely replace one simplistic rule with another.

A well-designed probability sample offers a formal design-based foundation for population inference, but it may itself suffer from coverage problems, nonresponse, measurement error, or inadequate sample size. A large non-probability dataset may contain valuable information and, with appropriate auxiliary data, assumptions, and statistical methods, may sometimes support useful population inference.

The point is not that smaller is better. The point is that size cannot substitute for understanding how the data came to represent the target population.

Can Weighting Rescue a Very Large Biased Sample?

Weighting, calibration, matching, propensity modeling, and data-integration methods can sometimes reduce differences between a sample and a target population.

Their success depends on the quality of external population information, measurement comparability, model specification, and whether the variables needed to address selection have actually been observed.

If selection depends strongly on an unmeasured characteristic related to the outcome, adjusting perfectly for age, sex, region, and education may still leave consequential bias.

Large sample size does not remove these assumptions. Indeed, when random error is already tiny, uncertainty about those assumptions may become more important than collecting another few thousand observations.

What Should You Examine Instead of Sample Size Alone?

Start with the target population and trace the path by which observations entered the dataset.

Who was covered? Who could be selected? Who encountered recruitment? Who volunteered or responded? Who remained in the analysis? What adjustments were applied? Which differences between the observed sample and target population are measurable, and which remain unknown?

These questions connect sample size with the broader issue of sampling bias. A million observations do not make those questions obsolete. They make answering them more consequential.

04 · A Practical Example

When 100,000 Responses Can Be Less Informative Than 2,000 Carefully Sampled Ones

Hypothetical Example

Estimating Generative AI Use Among University Students

Two researchers want to estimate the proportion of students in the same defined university population who use generative AI for academic work.

Study A An appropriate probability sample of 2,000 students is selected from a sufficiently complete enrollment frame. Recruitment and analysis follow the sample design, and consequential nonresponse is evaluated.
Study B An open online questionnaire receives 100,000 responses after being promoted through student technology communities, AI-related pages, and users who voluntarily share the survey.
Study B's numerical advantage The raw dataset contains 50 times as many observations, and conventional standard errors calculated as though the sample were simple random observations may be extremely small.
Study B's selection problem Interest in technology and AI may affect both exposure to recruitment and willingness to respond, while also being directly related to the outcome being estimated.
The inferential consequence For estimating prevalence in the defined university population, Study A may provide the stronger design-based foundation despite its much smaller participant count.
What cannot be concluded This does not prove that Study B's estimate is wrong or that every probability sample is superior. It shows that sample size alone cannot establish which estimate is more credible for the target population.
05 · What Researchers Often Get Wrong

Common Misconceptions About Large Samples and Bias

Misconception

Does Bias Average Out When the Sample Gets Large?

Random errors may average out under suitable conditions, but systematic selection does not necessarily do so. If the sampling process repeatedly favors the same type of participant, increasing the number selected through that process can preserve the systematic difference.

Misconception

Does a Tiny Margin of Error Mean the Estimate Must Be Accurate?

No. A conventional margin of error typically describes sampling uncertainty under particular assumptions. It does not automatically quantify unknown coverage, selection, nonresponse, measurement, or model error.

Misconception

If I Survey Most of the Population, Does Selection Stop Mattering?

Not necessarily. A high sampling fraction can greatly reduce sampling variability, but if the unobserved portion differs systematically on the quantity of interest, even a relatively small excluded group can matter. The mechanism determining inclusion still deserves examination.

Misconception

Can a Huge Convenience Sample Be Called Representative Because It Is Diverse?

No. Diversity and size do not by themselves establish representativeness. A sample can contain many demographic groups while systematically overrepresenting people who are easier or more motivated to participate.

Misconception

Does Statistical Significance Become More Trustworthy With a Huge Sample?

Larger samples can estimate associations more precisely under the model, but statistical significance does not correct selection bias, measurement problems, confounding, or other design limitations. It also does not establish substantive importance.

Misconception

Can Weighting Always Fix a Large Biased Dataset?

No. Adjustment can reduce some measurable differences under defensible assumptions, but it cannot guarantee correction for unobserved factors related to both selection and the outcome.

06 · What This Means for You

When Your Dataset Is Huge, Examine Selection More Carefully, Not Less

A simple decision framework

If the sample is large because recruitment was easy through one particular channel
Investigate which segments of the target population that channel disproportionately reaches or misses.
If conventional standard errors are extremely small
Separate random sampling uncertainty from uncertainty about coverage, selection, measurement, and model assumptions.
If the sampling frame excludes a meaningful population segment
Improve coverage where feasible rather than simply increasing the sample drawn from the covered portion.
If the data come from volunteers or opt-in participants
Examine how participation may relate to the outcomes and identify what auxiliary population information is available for adjustment or sensitivity analysis.
If additional observations barely change random precision
Invest resources in understanding data quality, selection, measurement, and target-population differences rather than assuming that more rows are automatically the best improvement.

A useful diagnostic question is: If I multiplied this sample by ten using exactly the same recruitment process, which methodological problem would actually disappear?

If the answer is only “the standard error would become smaller,” then any systematic selection problem would still need its own solution.

07 · A Quick Checklist

Before You Trust a Result Because the Sample Is Large

Look beyond the participant count:
Is the target population clearly defined independently of the dataset you happen to possess?
Does the sampling frame or data source adequately cover important parts of that population?
Could inclusion in the dataset be associated with the outcome or effect you are estimating?
Could volunteering, nonresponse, attrition, or platform participation systematically shape the observed sample?
Are narrow confidence intervals being interpreted only for the uncertainty they actually quantify?
Have you compared the sample with reliable target-population information on characteristics relevant to selection and the outcome where possible?
If statistical adjustments are used, have their assumptions and remaining limitations been made explicit?
Would the substantive conclusion remain persuasive if the sample size were not displayed prominently?
08 · Frequently Asked Questions

Questions About Large Samples and Bias

Can a large sample still be biased?

Yes. Increasing sample size can reduce random sampling variability while systematic differences between the observed sample and target population remain. A large sample can therefore be precise but biased.

Does bias decrease when sample size increases?

Not necessarily. Some forms of random error decrease with sample size, but systematic coverage, selection, nonresponse, or measurement biases may remain unless their underlying mechanisms are addressed.

What is the Big Data Paradox?

It describes how very large datasets can produce highly precise yet substantially misleading population estimates when data inclusion is systematically related to the quantity being measured. Increasing data quantity does not automatically compensate for defective representativeness.

Can a small sample be more reliable than a large sample?

For some inferential purposes, yes. A smaller sample produced by an appropriate probability design may provide a stronger basis for estimating a target population than a much larger but highly selective sample, although the smaller sample will generally have greater sampling variability.

Does a narrow confidence interval mean there is little bias?

No. A confidence interval quantifies uncertainty under particular assumptions and usually does not incorporate unknown systematic selection or measurement bias. Precision and bias should be assessed separately.

Can big data solve sampling problems?

Large datasets can provide enormous analytical value, but size alone does not solve coverage or selection problems. Their usefulness for population inference depends on how observations entered the data, what target population is intended, and what design or analytical adjustments are defensible.

Can weighting fix large-sample bias?

Weighting and related adjustment methods can reduce some differences between a sample and target population when suitable auxiliary information and assumptions are available. They cannot guarantee correction for important unmeasured selection mechanisms.

09 · The Bottom Line

A Huge Sample Can Give You the Wrong Answer With Remarkable Confidence

The Bottom Line

A large sample can substantially reduce random uncertainty while leaving systematic coverage, selection, nonresponse, or measurement bias intact, so sample size alone cannot establish that an estimate accurately represents the target population.

When the dataset becomes very large, pay at least as much attention to how observations entered it as to how many observations it contains. Once random error becomes small, understanding systematic error may matter far more than collecting another zero for the sample-size column.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes