Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How to Define Your Population and Choose the Right Sample

A defensible sample begins with a clearly defined population. Learn how to move from your research question to the people or units you actually recruit without claiming more than your design can support.

77
Define Your Population and Sample Guide 77 of 217
01 · The Question

Who Exactly Should Be in Your Study?

You have a research question, perhaps even a methodology and a data-collection instrument. Then comes a deceptively simple question: Who should you actually study?

This is where terms such as population, sample, target population, accessible population, and sampling frame begin to appear. They are related, but they do not describe the same thing. More importantly, choosing participants is not simply a matter of finding enough willing people. Your sample needs to make sense in relation to the question you are asking and the population about which you ultimately want to draw conclusions.

The challenge is therefore not merely deciding how many participants you need. It is establishing a defensible chain from research question → population → accessible participants or units → sampling approach → final sample.

02 · The Short Answer

Define the Population Before You Choose the Sample

In Brief

First define the population your research question concerns, then determine which members of that population you can realistically identify or reach, and finally choose a sampling approach that produces a sample appropriate for your research design and intended conclusions.

A sample is not automatically appropriate because it is large or convenient. What matters is whether the route from the population to the participants or units actually studied is methodologically defensible and whether your conclusions remain within the limits of that route.

03 · What You Need to Know

From the Population You Care About to the Sample You Actually Study

Start With the Research Question, Not With Whoever Is Available

Sampling decisions should begin with the phenomenon and population implied by your research question. Ask: Whom or what am I trying to understand?

If you want to estimate the prevalence of academic burnout among undergraduate nursing students at a university, your population concerns undergraduate nursing students within the scope you have specified. If you want to understand how experienced school principals make decisions during organizational crises, the relevant population is defined differently. If you are analyzing research articles rather than people, the population consists of documents or other units of analysis.

This distinction matters because a research population does not necessarily mean a population of human participants. Depending on the study, the units may be households, schools, organizations, patient records, publications, social-media posts, events, or other entities.

The population should therefore follow from what the study is trying to learn. Beginning instead with an easily available group and only afterward constructing a research question around that group reverses the logic of research design.

Define the Population Precisely Enough That Membership Is Clear

A statement such as “the population consists of teachers” is usually too broad to guide sampling. Which teachers? Where? Teaching at what level? During what period? With what characteristics relevant to the research question?

A useful population definition commonly specifies the characteristics that determine whether a unit belongs to the population. Depending on the study, these might include role, setting, geographic location, institutional affiliation, experience, exposure to a phenomenon, or a relevant time period.

For example:

Too vague: College students.

More useful: Undergraduate students enrolled in fully online degree programs at the university during the second semester of the academic year under study.

The second definition makes it much easier to determine who belongs, who does not, and what sources might be used to identify potential participants. The exact boundaries should be justified by the research question rather than added merely to make recruitment easier.

Separate the Population You Want to Understand From the Population You Can Reach

Researchers often discover that the population they conceptually want to study is broader than the group they can realistically access. That difference should be recognized rather than quietly ignored.

Your intended population and the population you can actually access may differ because of geography, institutional permissions, available records, recruitment channels, cost, time, or other practical constraints.

Suppose your research question concerns public-school teachers in an entire region, but you have permission to recruit only from five schools in one division. The five schools do not become equivalent to the regional population simply because they are the sites available to you. The distinction has consequences for sampling and, later, for the scope of your conclusions.

Population of interest The broader set of people, cases, or units the research question is intended to address.
Sample The subset of people, cases, or units from which the study actually obtains data.

The conceptual difference between a sample and the population it is intended to inform is fundamental. Studying a subset is often necessary, but conclusions about a larger population require a defensible basis for moving from what was observed in the sample to what is claimed about people or units beyond it.

Specify Who Is Eligible Before Recruitment Begins

Once the population is conceptually defined, translate that definition into operational eligibility rules. These rules determine who can enter the study and, when appropriate, who should not.

For example, a study of university instructors' experiences using generative AI in teaching might require participants to be currently teaching at least one higher-education course and to have used a generative AI application for a specified instructional purpose. Someone who has only read about generative AI would not necessarily belong in the same population if direct experience is central to the research question.

Your inclusion and exclusion criteria should make eligibility reproducible rather than subjective. At the same time, every restriction narrows the population represented by the study. Criteria should therefore have substantive or methodological justification rather than being added reflexively.

Determine How Members of the Population Can Actually Be Identified

Defining a population does not automatically give you a way to reach its members. You need some mechanism for identifying, approaching, or recruiting eligible units.

In survey research, this may involve a sampling frame: a list or other operational source from which potential respondents can be selected. The American Association for Public Opinion Research describes a sampling frame as the information that allows potential respondents to be contacted from a population. Depending on the study, this might be a student enrollment list, employee roster, membership registry, address database, telephone list, or panel.

The frame matters because it may not perfectly cover the population. A university email directory, for example, cannot represent eligible students who are missing from that directory. A professional-association membership list does not automatically represent every professional working in that field.

When a sampling frame does not adequately match the intended population, some eligible units may have little or no opportunity to enter the sample. This is a coverage problem, and increasing the number recruited from the same incomplete frame does not necessarily solve it.

Choose a Sampling Strategy Based on the Inference You Need

Only after clarifying the population, eligibility, and realistic source of participants should you decide how units will be selected.

The broad distinction is between probability and non-probability sampling. In probability sampling, selection is governed by a probability mechanism in which units have known selection probabilities under the design. In non-probability sampling, participants or units are selected without such known probabilities, for example through purposive recruitment, convenience access, quotas, volunteer participation, or network-based approaches.

Research need Sampling consideration Question to ask
Estimate characteristics of a defined population A probability-based design may be preferable when a suitable frame and resources are available Can eligible units be identified and selected through a defensible probability mechanism?
Understand a particular experience or phenomenon in depth Purposeful selection of information-rich cases may be appropriate Which participants can provide evidence directly relevant to the phenomenon?
Study a difficult-to-identify or difficult-to-reach group Network-based or other non-probability approaches may be necessary How can eligible participants realistically be located and recruited?
Conduct an exploratory or feasibility study under substantial constraints A practical sample may be defensible if its limitations are explicit What conclusions can this sample support, and which conclusions should I avoid?

The labels “probability” and “non-probability” are only the beginning. Specific designs include simple random, systematic, stratified, cluster, convenience, purposive, quota, snowball, and other approaches. The appropriate choice depends on the research objective, population structure, available frame, resources, analysis, and type of inference sought.

Do Not Choose Sample Size Before You Know What the Sample Is Supposed to Do

Researchers sometimes jump from “Who should I study?” directly to “How many respondents do I need?” A sample-size calculation cannot repair an ill-defined population or an unsuitable recruitment process.

For quantitative studies, the appropriate sample size and statistical power depend on the planned analysis, desired precision or power, assumptions about the expected effect or parameter, design features, and anticipated data loss or nonresponse. Different quantitative objectives therefore produce different sample-size requirements.

Qualitative research follows a different logic. Participant selection is usually tied to the phenomenon, analytic approach, heterogeneity of participants, richness of the data, and the informational needs of the study rather than a conventional statistical power calculation.

Watch Out

Do not treat a large sample as evidence that the sampling design is sound. Thousands of observations drawn through a systematically distorted recruitment process can still provide a poor basis for claims about the intended population.

Keep the Intended Conclusions in View Throughout the Sampling Process

The population and sample are connected to what you will eventually be allowed to say. If your evidence comes from a narrow accessible population or a highly selective sample, your conclusions should reflect that boundary.

This does not mean that every worthwhile study needs a statistically representative sample. Some research seeks population estimates, while other research aims to understand mechanisms, experiences, processes, cases, or theoretical patterns. Qualitative studies, case studies, experiments, and exploratory research may pursue forms of inference that differ substantially from those of a probability survey.

The important point is alignment: the claims you make should be compatible with how the sample was constructed and what the study was designed to accomplish.

04 · A Practical Example

From a Broad Research Idea to a Defensible Sample

Hypothetical Example

Studying Generative AI Use Among University Instructors

Suppose a researcher wants to estimate how university instructors at a particular university are using generative AI in their teaching during the current academic year. The university has several colleges and employs both full-time and part-time instructors.

1. Clarify the question The researcher wants to describe instructional use of generative AI among instructors currently teaching at the university, rather than among educators in general.
2. Define the population The population is defined as all full-time and part-time instructors teaching at least one course at the university during the specified academic period.
3. Establish eligibility Faculty members who are not teaching during the period are excluded because the question concerns current instructional practice.
4. Identify a sampling frame With institutional authorization, the researcher determines whether the official teaching roster contains all instructors meeting the population definition and investigates any meaningful omissions.
5. Choose the sampling approach Because the researcher wants population-level estimates and instructors differ across colleges and employment categories, a probability-based design that preserves relevant subgroups may be considered if the roster is sufficiently complete and recruitment is feasible.
6. Determine sample size The researcher calculates the required sample only after deciding what estimates or statistical analyses the study must support and accounts for the sampling design and expected nonresponse.
7. Recruit and document the achieved sample The researcher records who was eligible, how participants were selected and contacted, how many responded, and whether response patterns raise concerns about differences between respondents and the intended population.
8. Match conclusions to the evidence When reporting the findings, the researcher distinguishes the defined population, sampling frame, selected sample, and responding sample rather than treating them as interchangeable.

Notice that “How many instructors should I survey?” appears relatively late in this process. That is deliberate. A numerical sample-size target becomes meaningful only after the researcher knows who belongs to the population, how those people can be reached, and what the resulting data are intended to support.

05 · What Researchers Often Get Wrong

Common Mistakes When Defining a Population and Sample

Misconception

Is My Population Simply Everyone Who Could Answer My Questionnaire?

No. The population should be defined by the research question and study scope. People do not become members of the relevant population merely because they are available, willing, or capable of completing the instrument.

Misconception

Can I Define the Population After I Know Who Agreed to Participate?

That reverses the methodological sequence. Your population and eligibility rules should ordinarily be established before recruitment. Otherwise, the population definition can become a post hoc description of whoever happened to respond.

Misconception

Does a Large Sample Automatically Represent the Population?

No. Sample size and representativeness are different issues. A large sample can still systematically exclude important portions of the intended population. More observations generally reduce sampling variability under an appropriate design, but they do not automatically remove coverage, selection, or nonresponse problems.

Misconception

If Participants Meet the Inclusion Criteria, Is the Sample Automatically Appropriate?

Eligibility is only one part of the sampling problem. You must also consider where eligible participants came from, how they were selected or recruited, who had little or no opportunity to participate, and whether the resulting sample supports the intended analysis and conclusions.

Misconception

Do All Studies Need a Representative Sample?

No. The need for statistical representativeness depends on the purpose of the research and the type of inference being attempted. A study seeking population estimates has different sampling requirements from a qualitative inquiry seeking rich understanding of a particular experience. Calling every nonrepresentative sample defective ignores those methodological differences.

Misconception

Can Statistical Analysis Fix a Poorly Chosen Population?

Not in general. Sophisticated analysis cannot make observations answer a question about a fundamentally different population. Statistical adjustments may address some known imbalances under appropriate assumptions, but they do not erase every form of undercoverage, selection bias, or mismatch between the research question and the sampled population.

06 · What This Means for You

Build Your Sampling Plan as a Chain of Decisions

Before recruiting anyone, write down the complete path from your research question to your final sample. Each step should follow logically from the previous one. If one link is difficult to justify, that is where your design needs more attention.

A simple decision framework

If you cannot state exactly whom or what your research question concerns
Refine the population definition before making sampling decisions.
If the population you can access is narrower than the population you want to understand
Document the difference and reconsider whether your intended claims should also become narrower.
If you need estimates intended to describe a defined population
Examine whether a suitable sampling frame and probability-based selection process are feasible, while also planning for coverage and nonresponse problems.
If your objective requires participants with particular experiences, knowledge, or characteristics
Use a selection strategy consistent with that purpose and explain why those cases provide the evidence the study requires.
If practical constraints force you to use whoever is readily available
Treat convenience as a design limitation rather than evidence that the accessible group represents a broader population.
If you are deciding how many participants to recruit
Determine the required sample size using criteria appropriate to your methodology only after the population, sampling approach, and analytical purpose are clear.

A useful final test is to complete this sentence before data collection:

“We want to learn about ______, we can identify or recruit eligible units through ______, we will select them using ______ because ______, and this design allows us to make conclusions about ______ within the following limitations: ______.”

If you cannot complete that statement coherently, calculating a sample size is probably premature.

07 · A Quick Checklist

Before You Finalize Your Population and Sampling Plan

Before recruiting participants or selecting units, check:
Can you state clearly whom or what the research question is intended to describe, explain, compare, or understand?
Is your population defined precisely enough that you can determine whether a particular person, case, or unit belongs to it?
Have you distinguished the population you ultimately care about from the people or units you can realistically access?
Are your eligibility criteria justified by the research question rather than by convenience alone?
If you are using a sampling frame, have you examined who may be missing from it or incorrectly included?
Does your sampling method fit the type of evidence and inference your study requires?
Have you considered whether recruitment, refusal, or nonresponse could systematically change the composition of the achieved sample?
Are you determining sample size using principles appropriate to your methodology rather than relying on a universal number or percentage?
Will your eventual conclusions remain within what this population and sampling process can reasonably support?
08 · Frequently Asked Questions

Questions About Populations and Samples in Research

What is the difference between a population and a sample?

The population is the complete set of people, cases, or units defined as relevant to the research question, while the sample is the subset actually selected or observed. The conclusions you can draw beyond the sample depend partly on how that sample was obtained and on the study design.

Does the population have to be people?

No. A population can consist of people, organizations, schools, households, records, publications, events, online content, or other units of analysis. What constitutes the population depends on the research question.

Should I determine the population or sample size first?

Define the population and clarify the sampling design first. A sample-size calculation requires assumptions about what you are estimating or testing and, in quantitative research, often about the statistical and sampling design. In qualitative research, sample adequacy is assessed using different methodological considerations.

Can my accessible population be smaller than my target population?

Yes, and this is common. The important issue is to acknowledge the difference and consider whether limitations in access change the population your evidence can reasonably inform.

What if there is no complete list of my population?

A perfect sampling frame often does not exist. You may need to identify the best available frame, combine sources, use another recruitment strategy, or reconsider the sampling design. Whatever approach you use, document important gaps between the source of participants and the population you intend to study.

Can I use convenience sampling?

Sometimes. Convenience sampling may be defensible for particular exploratory, pilot, feasibility, educational, or otherwise constrained research purposes, but convenience does not by itself justify generalizing the findings to a broader population. The method should fit the purpose and its limitations should be explicit.

How specific should my population definition be?

Specific enough that another researcher could understand which units belong to the population and why. Include characteristics, setting, location, time period, or other boundaries when they are substantively relevant, but avoid restrictions that have no methodological justification.

Does a bigger sample solve sampling bias?

No. Increasing sample size can improve precision under suitable sampling and analytical conditions, but it does not automatically correct systematic undercoverage, self-selection, nonresponse, or other processes that make the observed sample differ meaningfully from the population relevant to the research question.

09 · The Bottom Line

Your Sample Should Follow From the Population, Not the Other Way Around

The Bottom Line

Define whom or what your research question concerns first, establish who is eligible and realistically reachable, and then choose a sampling strategy that fits the evidence and conclusions your study requires.

There is no universally correct sample. A defensible sampling plan is one in which the population, access, eligibility rules, selection process, sample size logic, and intended claims fit together, with important limitations made explicit rather than hidden behind the final number of participants.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes