03 · What You Need to Know
A Defensible Exclusion Must Fit the Question and Survive Its Consequences
Begin with the target of the research question
The strongest justification for excluding a population is often that the group is not part of the population the research question is designed to investigate.
Suppose a study examines how students experience the transition from secondary school to university-level academic writing. Restricting participation to first-year undergraduates can be methodologically coherent because transition into university is part of the phenomenon itself. Students in later years may have relevant experiences with academic writing, but they are no longer experiencing the same transition.
Now change the research question to "How do undergraduate students use generative AI for academic writing?" If the researchers still include only first-year students, they need to explain why evidence from that subgroup is appropriate for a question phrased around undergraduates generally. Otherwise, the population boundary and the intended inference no longer align.
A useful first test is therefore simple: Does the exclusion help define the population the question is actually about, or does it merely reduce the population available to answer a broader question?
Population exclusions can improve relevance and comparability
In participant-based research, eligibility criteria identify who can appropriately contribute evidence to the study. Methodological guidance describes inclusion criteria as characteristics defining the target population needed to answer the research question, while exclusion criteria can identify otherwise eligible people whose additional characteristics could interfere with participation, bias particular results, or create safety concerns.
For example, some clinical research may exclude participants with a comorbidity when that condition could substantially affect the outcome being studied. Intervention studies may exclude individuals for whom participation would create unacceptable risk. Longitudinal research may sometimes establish eligibility requirements connected to the ability to complete necessary follow-up procedures.
These examples do not establish a universal list of valid exclusions. They illustrate the underlying principle: an eligibility criterion should have a methodological or ethical function connected to the particular study.
Researchers should also consider the consequences of those criteria for external validity. Restricting the study population can make the sample more suitable for answering a particular question while simultaneously narrowing the population to which the findings can reasonably be generalized.
Participant safety can provide a strong justification for exclusion
In intervention research, some exclusions exist primarily because participation could expose particular individuals to unacceptable risk.
Clinical-trial eligibility criteria commonly consider characteristics such as health conditions, treatment history, disease status, and other factors relevant to whether participation is safe and appropriate. NIH guidance notes that inclusion and exclusion criteria help researchers identify appropriate participants and protect participant safety.
The logic extends beyond clinical trials, although the specific risks differ. A study involving psychologically distressing material, physically demanding tasks, inaccessible procedures, or particular privacy risks may require eligibility decisions related to participant welfare.
Ethical justification does not mean researchers should exclude populations merely because including them requires additional safeguards or accessibility measures. The ethical question is whether the exclusion itself is justified relative to the risks, benefits, scientific purpose, and available protections.
Exclusion can sometimes reduce unwanted heterogeneity
Researchers may restrict a population to reduce variation that would interfere with answering a focused question.
Imagine an educational intervention designed specifically for novice programmers. Including advanced computer science students may introduce substantial differences in prior expertise that are irrelevant to the intended population and could complicate interpretation of the intervention's effects.
Restricting participation to learners within a defined level of prior experience may therefore improve alignment between the intervention, population, and research question.
But greater homogeneity is not automatically better. If meaningful variation is part of the phenomenon the study is supposed to explain, removing it can make the study less informative. Eligibility criteria that are too narrow can also reduce the applicability of findings to broader populations. Methodological guidance on randomized trials therefore emphasizes balancing restrictions intended to improve study integrity against their consequences for generalizability.
Exclusion criteria and study delimitations operate at different levels
It helps to distinguish a broad population boundary from a specific eligibility rule.
Population delimitation
Defines the population the study deliberately concerns, such as first-year undergraduate students rather than all university students.
Exclusion criterion
Identifies a characteristic that makes an otherwise potentially eligible participant or case ineligible under the study protocol.
If the target population is first-year undergraduates, "not being a postgraduate student" usually does not need to become a separate exclusion criterion. The population has already been positively defined. Methodological guidance identifies using the same characteristic redundantly as both an inclusion and exclusion criterion as a common error.
Understanding what makes a boundary a delimitation can help prevent the eligibility section from becoming a list of everyone and everything outside the study.
Variables need justification for inclusion as well as exclusion
Questions about variable exclusion require somewhat different reasoning. A population determines who or what contributes evidence. Variables determine what characteristics, exposures, outcomes, relationships, or other quantities are measured or incorporated into an analysis.
Researchers sometimes assume that leaving out a variable requires justification while including additional variables is inherently safer. That is not always true.
Every variable should have a methodological role. Depending on the research design, a variable might represent an exposure, outcome, predictor, confounder, mediator, moderator, covariate, descriptive characteristic, or another construct required by the conceptual and analytical strategy.
A variable does not belong merely because previous studies have mentioned it. Nor does it belong simply because it is available in the dataset.
The useful question is: What role would this variable play in answering the research question?
Some variables can reasonably be excluded because they are outside the conceptual model
Suppose a study investigates the association between students' use of generative AI for academic writing and writing self-efficacy. The available dataset also contains satisfaction with campus food, commuting time, preferred social-media platform, and dozens of other characteristics.
Those variables may have interesting associations with student life. Their availability does not create a methodological obligation to analyze them.
Even variables more obviously connected with learning do not automatically belong. Motivation, anxiety, digital literacy, creativity, academic performance, and technology acceptance might all be relevant to the broad topic, but each needs a defensible role within the current research question and conceptual model.
Excluding theoretically unnecessary variables can reduce unfocused analysis and help maintain correspondence between the question and the evidence actually being interpreted.
But excluding a confounder can seriously distort an observational analysis
Variable exclusion becomes more consequential when the omitted variable is needed for the intended inference.
In observational research, a confounder is associated with the exposure and outcome in a way that can distort the exposure-outcome relationship if it is not appropriately addressed. Which variables require adjustment depends on the causal question and assumptions of the study; it cannot be decided simply by looking for statistical significance.
Suppose researchers examine whether use of an AI tutoring system is associated with academic performance. If prior academic achievement affects both students' likelihood of using the system and their subsequent performance, ignoring prior achievement could produce a misleading estimate of the relationship attributed to AI use.
The relevant question is not whether adding prior achievement makes the model more complicated. It is whether the intended interpretation requires accounting for it.
Watch Out
Do not exclude a variable merely because including it weakens, reverses, or makes a desired association statistically non-significant. Analytical decisions should follow the research question and defensible methodological reasoning rather than the direction of the result.
Including every possible covariate is not the solution either
If omitting important variables can create bias, it may seem safer to adjust for everything available. That strategy can also be problematic.
Variables can occupy different positions in the causal structure underlying a research question. Some may be confounders requiring adjustment, while others may be mediators, colliders, or consequences of the exposure. Adjusting indiscriminately can therefore change the quantity being estimated or introduce bias.
This is one reason covariate selection should be guided by substantive knowledge, a clearly defined estimand or analytical question, and an appropriate causal or statistical rationale rather than by automatic procedures alone.
For simpler descriptive or predictive studies, the precise analytical considerations differ, but the general principle remains: variable inclusion and exclusion should serve the stated purpose of the analysis.
Excluding a population can affect internal and external validity differently
Population restriction can have several methodological consequences, and they should not be collapsed into a single judgment that restriction is either "good" or "bad."
A narrower eligibility criterion might improve comparability or reduce a particular source of unwanted variation. At the same time, the resulting population may differ substantially from people who were excluded, limiting how readily the findings can be applied outside the study population.
Patino and Ferreira emphasize that researchers should evaluate how inclusion and exclusion decisions affect external validity. Their example illustrates that excluding participants with comorbidities can leave uncertainty about whether findings apply to people who have those comorbidities.
Restriction can also create selection-related problems under particular causal structures. Methodological work on selection bias shows that selecting a study or analytical sample based on characteristics related to exposures, outcomes, or their causes can sometimes distort associations rather than merely reduce generalizability.
That is why the consequences of an exclusion depend on how the selection mechanism relates to the question being estimated, not simply on how many people remain.
Convenience is a practical consideration, not a complete methodological justification
Researchers rarely have unlimited access to populations. A student researcher may have permission to recruit from one institution but not ten. A research team may have access to one administrative dataset but not another. Those realities matter.
However, accessibility should not be confused with conceptual relevance.
If the research question concerns students at one institution, access to that institution is coherent with the scope. If the research question claims to concern all university students nationally but the researchers include one conveniently available class and simply label everyone else "excluded," the terminology does not resolve the mismatch.
Research on registry populations similarly notes that convenience in determining the accessible population can reduce representativeness when readily enrolled individuals differ meaningfully from the population of interest.
When access forces a narrower study, the better response is often to narrow the research question and intended claims as well.
The exclusion must be judged against the claims you intend to make
The same population restriction can be defensible for one claim and inadequate for another.
Consider a study restricted to first-year nursing students at one university.
If the claim is about the experiences of first-year nursing students in that institutional context, the boundary may be entirely appropriate.
If the conclusion becomes "university students prefer AI-assisted learning," the problem is not necessarily that first-year nursing students were excluded from nothing. The problem is that the conclusion extends beyond the population represented by the evidence.
This is why eligibility decisions affect external validity. Researchers should ask not only whether the selected participants can answer the question but also what population the findings are ultimately intended to inform.
A transparent exclusion is easier to evaluate than an invisible one
Methodological justification requires enough reporting for readers to understand what was excluded and why.
For participant research, eligibility criteria should normally be established during study design rather than invented after researchers see who produces convenient results. In evidence synthesis, authoritative guidance similarly recommends prespecifying eligibility criteria in the protocol and documenting and justifying subsequent changes.
The same logic applies more broadly. Researchers should be particularly cautious about exclusions introduced after observing the data or results.
If an exclusion changes during the study, report what changed and why. If cases are removed during analysis, explain the analytical criterion. If a variable planned in the protocol cannot be analyzed, state the reason rather than silently omitting it.
Methodological justification does not mean pretending the exclusion has no cost
A decision can be justified and still involve a trade-off.
Restricting an intervention study to participants who can safely receive the intervention may be ethically necessary while limiting applicability to people with excluded conditions. Restricting a qualitative study to one stakeholder group may permit greater depth while leaving other perspectives unexamined. Excluding a variable because it lies outside the conceptual model may sharpen the study while leaving another plausible explanatory pathway for future investigation.
Good methodological reasoning acknowledges those consequences instead of treating "justified" as synonymous with "consequence-free."
This is the difference between defending a boundary and denying that a boundary exists.