03 · What You Need to Know
From the Population You Care About to the Sample You Actually Study
Start With the Research Question, Not With Whoever Is Available
Sampling decisions should begin with the phenomenon and population implied by your research question. Ask: Whom or what am I trying to understand?
If you want to estimate the prevalence of academic burnout among undergraduate nursing students at a university, your population concerns undergraduate nursing students within the scope you have specified. If you want to understand how experienced school principals make decisions during organizational crises, the relevant population is defined differently. If you are analyzing research articles rather than people, the population consists of documents or other units of analysis.
This distinction matters because a research population does not necessarily mean a population of human participants. Depending on the study, the units may be households, schools, organizations, patient records, publications, social-media posts, events, or other entities.
The population should therefore follow from what the study is trying to learn. Beginning instead with an easily available group and only afterward constructing a research question around that group reverses the logic of research design.
Define the Population Precisely Enough That Membership Is Clear
A statement such as “the population consists of teachers” is usually too broad to guide sampling. Which teachers? Where? Teaching at what level? During what period? With what characteristics relevant to the research question?
A useful population definition commonly specifies the characteristics that determine whether a unit belongs to the population. Depending on the study, these might include role, setting, geographic location, institutional affiliation, experience, exposure to a phenomenon, or a relevant time period.
For example:
Too vague: College students.
More useful: Undergraduate students enrolled in fully online degree programs at the university during the second semester of the academic year under study.
The second definition makes it much easier to determine who belongs, who does not, and what sources might be used to identify potential participants. The exact boundaries should be justified by the research question rather than added merely to make recruitment easier.
Separate the Population You Want to Understand From the Population You Can Reach
Researchers often discover that the population they conceptually want to study is broader than the group they can realistically access. That difference should be recognized rather than quietly ignored.
Your intended population and the population you can actually access may differ because of geography, institutional permissions, available records, recruitment channels, cost, time, or other practical constraints.
Suppose your research question concerns public-school teachers in an entire region, but you have permission to recruit only from five schools in one division. The five schools do not become equivalent to the regional population simply because they are the sites available to you. The distinction has consequences for sampling and, later, for the scope of your conclusions.
Population of interest
The broader set of people, cases, or units the research question is intended to address.
Sample
The subset of people, cases, or units from which the study actually obtains data.
The conceptual difference between a sample and the population it is intended to inform is fundamental. Studying a subset is often necessary, but conclusions about a larger population require a defensible basis for moving from what was observed in the sample to what is claimed about people or units beyond it.
Specify Who Is Eligible Before Recruitment Begins
Once the population is conceptually defined, translate that definition into operational eligibility rules. These rules determine who can enter the study and, when appropriate, who should not.
For example, a study of university instructors' experiences using generative AI in teaching might require participants to be currently teaching at least one higher-education course and to have used a generative AI application for a specified instructional purpose. Someone who has only read about generative AI would not necessarily belong in the same population if direct experience is central to the research question.
Your inclusion and exclusion criteria should make eligibility reproducible rather than subjective. At the same time, every restriction narrows the population represented by the study. Criteria should therefore have substantive or methodological justification rather than being added reflexively.
Determine How Members of the Population Can Actually Be Identified
Defining a population does not automatically give you a way to reach its members. You need some mechanism for identifying, approaching, or recruiting eligible units.
In survey research, this may involve a sampling frame: a list or other operational source from which potential respondents can be selected. The American Association for Public Opinion Research describes a sampling frame as the information that allows potential respondents to be contacted from a population. Depending on the study, this might be a student enrollment list, employee roster, membership registry, address database, telephone list, or panel.
The frame matters because it may not perfectly cover the population. A university email directory, for example, cannot represent eligible students who are missing from that directory. A professional-association membership list does not automatically represent every professional working in that field.
When a sampling frame does not adequately match the intended population, some eligible units may have little or no opportunity to enter the sample. This is a coverage problem, and increasing the number recruited from the same incomplete frame does not necessarily solve it.
Choose a Sampling Strategy Based on the Inference You Need
Only after clarifying the population, eligibility, and realistic source of participants should you decide how units will be selected.
The broad distinction is between probability and non-probability sampling. In probability sampling, selection is governed by a probability mechanism in which units have known selection probabilities under the design. In non-probability sampling, participants or units are selected without such known probabilities, for example through purposive recruitment, convenience access, quotas, volunteer participation, or network-based approaches.
| Research need |
Sampling consideration |
Question to ask |
| Estimate characteristics of a defined population |
A probability-based design may be preferable when a suitable frame and resources are available |
Can eligible units be identified and selected through a defensible probability mechanism? |
| Understand a particular experience or phenomenon in depth |
Purposeful selection of information-rich cases may be appropriate |
Which participants can provide evidence directly relevant to the phenomenon? |
| Study a difficult-to-identify or difficult-to-reach group |
Network-based or other non-probability approaches may be necessary |
How can eligible participants realistically be located and recruited? |
| Conduct an exploratory or feasibility study under substantial constraints |
A practical sample may be defensible if its limitations are explicit |
What conclusions can this sample support, and which conclusions should I avoid? |
The labels “probability” and “non-probability” are only the beginning. Specific designs include simple random, systematic, stratified, cluster, convenience, purposive, quota, snowball, and other approaches. The appropriate choice depends on the research objective, population structure, available frame, resources, analysis, and type of inference sought.
Do Not Choose Sample Size Before You Know What the Sample Is Supposed to Do
Researchers sometimes jump from “Who should I study?” directly to “How many respondents do I need?” A sample-size calculation cannot repair an ill-defined population or an unsuitable recruitment process.
For quantitative studies, the appropriate sample size and statistical power depend on the planned analysis, desired precision or power, assumptions about the expected effect or parameter, design features, and anticipated data loss or nonresponse. Different quantitative objectives therefore produce different sample-size requirements.
Qualitative research follows a different logic. Participant selection is usually tied to the phenomenon, analytic approach, heterogeneity of participants, richness of the data, and the informational needs of the study rather than a conventional statistical power calculation.
Watch Out
Do not treat a large sample as evidence that the sampling design is sound. Thousands of observations drawn through a systematically distorted recruitment process can still provide a poor basis for claims about the intended population.
Keep the Intended Conclusions in View Throughout the Sampling Process
The population and sample are connected to what you will eventually be allowed to say. If your evidence comes from a narrow accessible population or a highly selective sample, your conclusions should reflect that boundary.
This does not mean that every worthwhile study needs a statistically representative sample. Some research seeks population estimates, while other research aims to understand mechanisms, experiences, processes, cases, or theoretical patterns. Qualitative studies, case studies, experiments, and exploratory research may pursue forms of inference that differ substantially from those of a probability survey.
The important point is alignment: the claims you make should be compatible with how the sample was constructed and what the study was designed to accomplish.