03 · What You Need to Know
Sample Quality Is About Fit, Selection, and Information
Start with the population or phenomenon the study is trying to understand
You cannot decide whether a sample is appropriate until you know what it is supposed to represent or illuminate.
Suppose researchers want to estimate the prevalence of generative AI use among undergraduate students in the Philippines. Their target population might be all currently enrolled undergraduates in the country.
Now suppose they collect responses from students in one information technology program at one private university.
Those students may provide perfectly valid information about themselves. The problem arises if the evidence is treated as though it directly represents the much broader national undergraduate population.
Begin by distinguishing:
| Term |
Practical meaning |
| Target population |
The broader population or set of units about which the researchers ultimately want to make conclusions |
| Sampling frame |
The operational source or list from which units can actually be selected, when such a frame exists |
| Eligible population |
The units that satisfy the study's inclusion and exclusion criteria |
| Invited or approached sample |
The units researchers attempted to recruit or include |
| Participating sample |
The units that actually entered or contributed data to the study |
| Analytical sample |
The units ultimately included in a particular analysis |
Those groups can differ substantially. Critical appraisal requires noticing where those differences occur.
Ask whether the sample contains the units required by the research question
The first sampling question is not numerical. It is substantive.
If the study asks about novice teachers, did it actually study novice teachers? If the claim concerns adolescents, what ages were included? If researchers want to understand rural schools, how were rural settings defined and represented? If the study concerns organizations rather than individuals, were organizations sampled appropriately?
A large sample of the wrong population does not become appropriate through size.
This is closely connected to whether the overall study design answers the research question. Sampling is one of the mechanisms connecting the question to the evidence.
Do not assume that “random sample” and “random assignment” mean the same thing
These terms are frequently confused, but they solve different methodological problems.
Random sampling
Concerns how units are selected from a population and can support population inference when appropriately implemented.
Random assignment
Concerns how enrolled participants are allocated to study conditions and can strengthen causal comparisons between those conditions.
A randomized controlled trial can randomly assign participants without having randomly sampled them from the wider population.
For example, 100 volunteers from one university may be randomly allocated to an intervention or control condition. Random assignment can strengthen the internal comparison between those groups. It does not automatically make the 100 volunteers representative of all university students.
Conversely, a probability sample from a population can provide strong descriptive evidence about that population without involving any experimental assignment.
Keep the inferential roles separate.
Ask how participants or cases entered the study
The recruitment pathway can reveal important selection processes.
Did researchers use probability sampling? Recruit volunteers through advertisements? Invite all eligible members of an institution? Select intact classrooms? Recruit patients from one clinic? Use administrative records? Draw cases from a registry? Sample social-media users who responded to an open link?
None of these strategies is automatically unacceptable.
The important question is whether the mechanism of entry could produce systematic differences relevant to the outcome or claim.
For example, an online voluntary survey about AI use may disproportionately attract students who are unusually interested in AI, strongly supportive of it, strongly opposed to it, or simply more active online. If those characteristics are related to the outcomes being estimated, the sample may differ meaningfully from the population the researchers intend to describe.
Convenience sampling is not automatically fatal
Convenience samples are common because researchers often have limited access to participants, institutions, records, or specialized populations.
A convenience sample can be useful for exploratory research, methodological development, experiments focused primarily on internal comparisons, qualitative inquiry, pilot studies, and other purposes where population representativeness is not the sole objective.
The problem arises when conclusions travel much farther than the sampling strategy permits.
Defensible conclusion
Among students participating in this study, X was associated with Y.
Potential overreach
University students generally show the same relationship, despite recruitment from one narrow convenience sample.
Do not reject a study merely because you see the phrase "convenience sample." Ask what inference depends on that sample being representative.
Probability sampling helps, but implementation still matters
Probability sampling gives population units a known mechanism of selection and can support statistical inference to a defined population when implemented appropriately.
But simply naming a probability-sampling method does not guarantee a representative final dataset.
A carefully drawn probability sample can still suffer from substantial nonresponse. The sampling frame may omit parts of the population. Some selected units may be unreachable. Weighting procedures may be necessary. Clustered or stratified designs may require corresponding analytical treatment.
So ask not only:
How was the sample selected?
Also ask:
Who actually provided the data?
Response rate alone does not tell you whether nonresponse bias is serious
A low response rate can be concerning, but the percentage alone does not determine the amount of bias.
Nonresponse bias depends on whether people who respond differ from people who do not respond in ways relevant to the quantities being estimated.
Imagine two surveys, each with a 50% response rate.
In the first, responders and nonresponders appear similar on characteristics strongly related to the outcome. In the second, nearly all students with low academic engagement fail to respond to a survey estimating engagement.
The response rate is identical. The potential inferential problem is not.
Look for information about recruitment, response patterns, comparisons between responders and nonresponders where available, weighting or adjustment procedures, and sensitivity analyses.
Inclusion and exclusion criteria shape the population you actually studied
Eligibility criteria are not administrative details. They define who can contribute evidence.
Researchers may exclude participants because of age, diagnosis, prior exposure, language, comorbidities, course enrollment, missing baseline information, technical requirements, or other considerations.
Those exclusions may be methodologically justified. They can improve interpretability by creating a more clearly defined population or reducing factors that would complicate the study.
They can also narrow applicability.
If a clinical trial excludes older adults, people with common comorbidities, and patients taking several medications, its results may be less directly applicable to real-world patients who frequently have those characteristics.
If an educational intervention study includes only students with reliable personal devices and high-speed internet, the findings may not tell you as much about implementation among students with limited digital access.
Ask whether the exclusions create a population meaningfully different from the one to which the conclusion is later applied.
Look at who was excluded after recruitment too
Selection does not end when participants enter the study.
Researchers may exclude observations during data cleaning or analysis because of missing values, incomplete responses, protocol deviations, outliers, failed attention checks, insufficient exposure, invalid measurements, or other criteria.
Some exclusions are necessary and defensible. Others can change the analytical sample in ways that matter.
Ask:
- Were exclusion rules defined clearly?
- Were they determined before or after seeing outcomes?
- How many observations were removed?
- Were exclusions similar across comparison groups?
- Could excluded participants differ systematically from those retained?
- Would alternative reasonable inclusion rules change the result?
The sample you evaluate is ultimately the sample that produced the reported estimate, not merely the number initially recruited.
Attrition can change an initially appropriate sample
Longitudinal and intervention studies may begin with a well-defined sample and gradually lose participants.
Attrition is especially concerning when loss is substantial or differs systematically according to group, exposure, outcome, or characteristics related to the outcome.
Suppose an intervention begins with 200 participants evenly divided between two groups. At follow-up, 95 control participants remain but only 55 intervention participants do because participants finding the intervention difficult were more likely to leave.
The remaining intervention group may no longer provide an unbiased picture of everyone originally assigned to that condition.
Do not merely compare the starting sample sizes. Trace the sample through the study.
Eligible Who could enter the study?
Recruited Who actually enrolled or contributed initial data?
Retained Who remained through the relevant follow-up?
Analyzed Whose data ultimately produced the reported result?
Each transition can introduce selection.
Sample size is about information, not merely headcount
Researchers often ask whether a sample is "big enough." There is no universally adequate number.
The information required depends on the research question and design.
For quantitative studies, relevant considerations can include expected effect sizes, outcome variability, desired precision, significance level, statistical power, number of parameters, event frequency, clustering, repeated observations, attrition, and the analytical model.
A sample of 100 may be ample for one analysis and inadequate for another.
For example, 1,000 participants might sound large. If only 20 experience the outcome required for a complex predictive model, the effective information available for that part of the analysis may still be limited.
This is why the question whether a small sample automatically makes a study weak cannot be answered from the number alone.
Look for a sample-size justification when the design calls for one
In many quantitative studies, authors should explain how the planned sample size was determined.
For randomized trials, CONSORT 2025 includes sample-size determination among the information that should be reported, including assumptions used in the calculation. Similar expectations appear across many study-specific reporting guidelines.
A sample-size calculation is not a magical certificate that the sample is adequate. Its usefulness depends on the assumptions entered into it, such as expected effect size, variability, attrition, or event rates.
Ask whether those assumptions are plausible and whether the final analyzed sample resembles the sample for which the calculation was performed.
Also distinguish between planning for statistical power and planning for estimation precision. In some studies, the more relevant question is whether the confidence interval will be narrow enough to answer the substantive question rather than whether a null-hypothesis test crosses a conventional threshold.
A non-significant result from a small sample may be inconclusive rather than negative
Suppose a study with 25 participants per group reports no statistically significant difference and concludes that the intervention has no effect.
Look at the estimate and confidence interval.
If the interval is wide enough to include both a meaningful benefit and a meaningful harm, the study may not have established absence of an effect. It may simply provide insufficiently precise evidence.
Sample size therefore affects what a null result can rule out.
Do not use p >.05 as evidence that the sample was adequate or that the groups are equivalent.
A huge sample does not guarantee representativeness
This is one of the most persistent sampling misconceptions.
Imagine an online survey receives 100,000 responses. That number can produce extremely precise estimates for the responding sample.
If participation was strongly self-selected and the respondents differ systematically from the target population, the estimate can still be biased.
A smaller probability sample may provide more defensible population inference than a vastly larger self-selected sample.
Watch Out
Large samples reduce some forms of random sampling uncertainty. They do not automatically eliminate selection bias, coverage problems, nonresponse bias, measurement error, or mismatch between the sample and the population of interest.
This is why a large sample should not be treated as an overall quality score.
Representativeness is always representativeness of something
Calling a sample "representative" is incomplete unless you specify the population.
A sample might approximate the demographic composition of one university while being unrepresentative of the national university population. A national sample might represent adults generally but contain too few members of a particular subgroup for precise subgroup estimates.
Ask:
Representative of which defined population, with respect to which characteristics relevant to this inference?
No sample can mirror a population perfectly on every possible variable. The relevant issue is whether differences between the sample and target population could materially affect the conclusion.
Representativeness is not equally important for every research question
External validity matters differently across studies.
If your goal is to estimate national prevalence, population sampling is central.
If your goal is to test whether a tightly controlled intervention can produce an effect under specified conditions, internal validity may initially receive more attention than population representativeness.
If your goal is to understand a rare or specialized experience qualitatively, purposeful selection of information-rich participants may be more appropriate than statistical representativeness.
Do not ask whether a sample is representative as though every study has the same inferential objective.
Generalizability depends on more than demographics
Researchers sometimes compare a sample and population on age and sex and conclude that the sample is representative.
Those characteristics may matter, but generalizability can depend on much more.
In educational research, institutional selectivity, academic program, prior achievement, technology access, instructional culture, socioeconomic context, language, and digital experience may influence whether an intervention transfers.
In clinical research, disease severity, comorbidities, treatment setting, prior therapy, and healthcare access may matter.
The characteristics that deserve attention are those that could modify the relationship or effect you intend to generalize.
Subgroup analyses need adequate information within the subgroup
A study can have a large overall sample and still provide weak evidence for particular subgroup claims.
Suppose a study includes 5,000 participants but only 37 belong to the subgroup for which the authors make a prominent conclusion.
The overall sample size does not determine the precision or stability of that subgroup estimate.
Check how many observations and events actually contribute to the comparison you care about.
Also be cautious when many subgroup analyses are conducted. A dramatic finding in one small subgroup may be exploratory, imprecise, or vulnerable to chance variation.
Clustered samples change the amount of independent information
Participants are often nested within classrooms, schools, hospitals, organizations, families, regions, or other clusters.
Two hundred students from two classrooms do not necessarily provide the same information as 200 students independently sampled from 200 different contexts. Students within the same classroom may resemble one another because they share instructors, curriculum, environment, or other influences.
The effective statistical information therefore depends partly on the number and structure of clusters, not only the total number of individual observations.
The analysis should account for that structure where relevant. This is one point at which sample evaluation connects directly to whether the analysis matches the design.
For qualitative research, “representative sample” may be the wrong standard
Qualitative research often uses purposive, theoretical, criterion-based, maximum-variation, snowball, or other non-probability approaches because the goal is not necessarily statistical estimation of a population parameter.
A qualitative researcher may deliberately select participants who have direct experience of the phenomenon, represent contrasting cases, or can illuminate particular aspects of the research question.
Judging such a sample solely by whether it is statistically representative misunderstands the design.
Instead ask:
- Does the sampling strategy fit the qualitative methodology and research question?
- Were participants capable of providing relevant information about the phenomenon?
- Is there sufficient variation or depth for the analytical purpose?
- Are important sampling decisions explained?
- Does the interpretation remain appropriately bounded by the participants and context studied?
This distinction becomes central when critically evaluating qualitative research on its own methodological terms.
Qualitative sample adequacy is not determined by a universal number
A qualitative study with 12 participants is not automatically weaker than one with 50.
Sample adequacy depends on the research question, methodological approach, specificity of the sample, richness of the data, analytical strategy, diversity sought, and the conceptual or interpretive goals of the study.
Some approaches discuss saturation, although what saturation means and how it should be assessed varies. Malterud and colleagues proposed the concept of information power, arguing that the more information a sample holds relevant to the study aim, the fewer participants may be needed. Their model considers factors such as the specificity of the sample, use of established theory, quality of dialogue, and analytical strategy.
The broader lesson is useful: do not import a quantitative power-analysis mindset into qualitative sampling without considering what the qualitative design is trying to accomplish.
Sampling in mixed-methods research may involve more than one sample
A mixed-methods study may use one sample for a survey, another subset for interviews, and perhaps additional cases for observations or follow-up.
Each component should be evaluated according to its own purpose.
You also need to ask how the samples relate. Were interview participants purposefully selected from survey respondents to explain particular quantitative patterns? Were the quantitative and qualitative samples drawn from unrelated populations even though the authors later integrate their findings?
Sampling adequacy in mixed-methods research therefore includes the relationship between samples as well as the quality of each one individually.
Samples can consist of more than people
Research samples may include schools, countries, organizations, documents, social-media posts, journal articles, images, biological specimens, administrative records, websites, events, or other units.
The same core questions apply:
What is the population or universe of interest? How were units selected? Which units could never enter the sample? Which were excluded? Does the final sample support the intended inference?
For example, a content analysis claiming to characterize "research articles about AI in education" needs a defensible process for defining and locating the relevant corpus. Sampling only papers from one database or only English-language publications may create boundaries that need to be acknowledged.
Secondary datasets inherit the sampling decisions of the original data collection
Large public or administrative datasets can create the impression that sampling is no longer an issue because the researcher did not recruit participants directly.
But someone or some system determined which observations entered the dataset.
Ask how the original data were generated, who is included, who is absent, what eligibility rules applied, whether records are complete, and whether the dataset was created for the research purpose now being imposed on it.
A dataset can contain millions of observations and still systematically exclude populations relevant to your question.
Ask who is missing
This is one of the simplest and most productive sampling questions.
When you read the participant characteristics, do not look only at who is there.
Ask:
Who could reasonably belong to the target population but had little or no chance of appearing in this sample?
Perhaps the study recruited through smartphones and therefore underrepresented people with limited digital access. Perhaps only English-language questionnaires were offered. Maybe participation required attending an optional workshop during working hours. Perhaps a hospital-based sample omits people who never access healthcare.
The missing group matters when its absence could change the finding or restrict the conclusion.
Ask whether the sample changed between the research question and the conclusion
Wording can gradually broaden as a paper progresses.
The methods may accurately describe "students enrolled in three introductory psychology courses at University X." The discussion may begin referring to "undergraduate students." The conclusion may eventually say "young adults."
Notice that expansion.
Actual sample Students in three courses at one institution.
Immediate evidence What was observed among those participants under the study conditions.
Target inference The broader population the authors want the result to inform.
Critical question What evidence justifies moving from the actual sample to that broader population?
Sometimes the extension is reasonable. Sometimes it requires considerable caution. What matters is that the inferential step is visible rather than automatic.
An inappropriate sample may narrow the conclusion rather than destroy the study
Suppose a study claims that an intervention improves writing performance among university students, but all participants are first-year engineering students at one institution.
The sample may be too narrow for the broadest claim.
But the study may still provide meaningful evidence that the intervention improved performance among the students actually studied.
This illustrates the distinction between an ordinary limitation and a more fundamental flaw. The sampling problem may restrict external validity without necessarily invalidating the internal comparison.
Ask whether the sampling issue merely narrows the inference or undermines the central comparison itself.