01 · The Question
Are You Interested in Your Participants or the Population Behind Them?
You survey 500 university students and find that 62% report using generative AI for academic work. What have you learned?
You certainly learned something about the students who responded. But can you say that 62% of all students at the university use generative AI? What about all university students in the country?
This is the practical difference between a sample and a population. Researchers usually observe a limited set of people, cases, records, or other units, yet their questions often concern a larger group. Understanding how those two groups relate is central to sampling, statistical inference, generalizability, and the claims you eventually make.
03 · What You Need to Know
Understanding the Relationship Between a Sample and a Population
What Is a Population in Research?
A population is the complete set of units to which a research question refers within its defined boundaries. Those units need not be people.
A population might consist of all undergraduate students enrolled at a university during a particular semester, all public secondary schools in a province, all articles published by a specified set of journals during a defined period, or all transactions meeting particular criteria in a database.
The phrase “entire population” therefore does not mean everyone in society. It means every unit belonging to the population as you have defined it.
That definition matters. “University students,” “university students in Manila,” and “first-year students enrolled in private universities in Manila during a specified academic year” describe different populations. A defensible study should make clear which one its research question actually concerns.
What Is a Sample?
A sample is a subset of units drawn, selected, recruited, or otherwise obtained from a population or an operationally reachable portion of it for observation or analysis.
If a university has 20,000 eligible students and your study obtains usable survey responses from 800 of them, those 800 respondents constitute the achieved sample. The population and sample are connected, but they are not interchangeable.
Population
The complete set of units defined as relevant to the research question.
Sample
The subset of units from which the study actually obtains evidence.
Why Do Researchers Study Samples Instead of Entire Populations?
Sometimes researchers can study every unit in a defined population. This is generally described as a census of that population. For a small organization, classroom, archive, or finite collection of records, complete enumeration may be feasible.
For many research questions, however, studying every unit would require too much time, money, access, or effort. Some populations are extremely large. Others change continuously or cannot be completely identified. Data collection itself may also be burdensome or destructive.
Sampling makes it possible to obtain evidence about a population without observing every member. Major statistical agencies routinely use samples for precisely this reason. The U.S. Census Bureau, for example, uses sample surveys to estimate population characteristics concerning employment, income, health insurance, educational attainment, business activity, and many other phenomena.
The methodological challenge is not merely obtaining fewer observations. It is designing the process so that the sample can answer the research question with an appropriate level of uncertainty and within clearly stated limits.
Population Values and Sample Estimates Are Not the Same Thing
In quantitative research, a numerical characteristic of a population is commonly called a parameter, while a numerical quantity calculated from a sample is a statistic.
Suppose the true proportion of all students in a defined university population who use generative AI is 0.58. That population proportion is a parameter. If 500 sampled students produce an observed proportion of 0.61, the sample proportion is a statistic that may be used to estimate the population parameter.
Parameter
A numerical characteristic of the defined population.
Statistic
A numerical quantity calculated from observed sample data.
In real research, the population parameter is often unknown. That is why the sample is collected in the first place.
Why Would Another Sample Give a Different Answer?
If you repeatedly drew samples from the same population using the same probability sampling design, you would not expect every sample to contain exactly the same units or produce exactly the same estimate. This sample-to-sample variation is the basis of sampling variability.
For probability samples, statistical methods can quantify uncertainty arising from the sampling process through measures such as standard errors and confidence intervals, provided the analysis properly reflects the sample design.
This is why an estimate from a sample should not ordinarily be interpreted as though it were the exact population value. The U.S. Census Bureau distinguishes sample estimates from the values that would be obtained under complete enumeration and reports measures of sampling uncertainty for its sample-based estimates.
A Bigger Sample Usually Improves Precision, but That Is Not the Whole Story
Under appropriate sampling conditions, larger samples generally provide more precise estimates than smaller samples. Sampling error tends to decrease as sample size increases.
But precision is not the same as validity or representativeness. If your recruitment process systematically misses an important part of the population, collecting thousands more observations through the same process can produce an extremely precise estimate for the wrong group.
This is why asking whether a bigger sample automatically makes a study better requires more than looking at the final participant count. Sample size, coverage, selection, measurement, nonresponse, and study design address different sources of uncertainty and error.
Can You Generalize From the Sample to the Population?
Sometimes, but not automatically.
If the objective is to estimate population characteristics using survey sampling, how units were selected matters substantially. A well-designed probability sample provides a formal basis for design-based inference because selection probabilities are known under the sampling design. Complex probability samples may also require weights and design-aware variance estimation.
A probability or non-probability sampling strategy should therefore be chosen with the intended inference in mind rather than treated merely as a procedural label in the methodology section.
Even probability sampling does not eliminate every problem. The sampling frame may omit parts of the population, selected participants may not respond, measurements may be inaccurate, and implementation may depart from the planned design.
Conversely, not every study is trying to estimate a population percentage or mean. Experiments, qualitative studies, case studies, and other designs may support different kinds of inference. The relationship between sample and population should therefore be interpreted in the context of the research question and design rather than through a single universal rule.
Representative of Which Population?
Researchers sometimes describe a sample simply as “representative,” as though representativeness were a property that could exist without a reference population.
It cannot. A sample can resemble or support inference to one population while failing to do so for another. A sample drawn from students at one institution might adequately support some claims about that institution under an appropriate design while providing a much weaker basis for claims about university students nationwide.
Before asking whether you have a representative sample, specify the population to which the description refers.
Your Final Sample May Be Several Steps Removed From the Population
In practice, researchers often move through several progressively narrower groups rather than directly from population to sample.
Target population The population the study ultimately seeks to inform.
Accessible population The portion that can realistically be reached under the study conditions.
Sampling frame or recruitment source The operational mechanism through which potential units are identified, when applicable.
Selected or invited sample The units chosen or approached for participation.
Achieved sample The units from which usable observations are ultimately obtained.
The distinction between your target and accessible populations becomes especially important when practical access narrows the study long before individual participants are selected.
Attrition or nonresponse can narrow it again. If people who respond differ systematically from those who do not, the achieved sample may no longer reflect even the group initially selected. This is why sample quality cannot be judged from the final number alone.