03 · What You Need to Know
Exclusion Should Follow From the Research Question, Not From Convenience Alone
Start with the evidence your research question actually requires
The easiest way to decide what can be left out is to begin with what must remain.
Take the research question apart. What population, phenomenon, constructs, comparisons, setting, period, or evidence must be represented for the question to be answered? Which elements are central to the claim you eventually want to make?
Suppose the question concerns how first-year university students experience the transition from secondary-school writing to AI-assisted academic writing. First-year status is not incidental. Removing that population characteristic changes the phenomenon being investigated. Likewise, removing academic writing and replacing it with general AI use would create a different inquiry.
By contrast, adding postgraduate students, faculty attitudes, institutional AI policies, student grades, and several unrelated psychological variables might broaden the project without helping answer that particular question.
A useful starting rule is therefore: protect what the question logically requires before deciding what the project can afford to exclude.
Relevance to the topic is not enough for inclusion
One of the reasons research projects expand so easily is that many things are genuinely related to the topic.
If you study academic performance, motivation is relevant. So are prior achievement, socioeconomic conditions, self-efficacy, attendance, teaching practices, cognitive ability, learning strategies, and many other factors. Their relevance does not mean that one study must measure all of them.
The inclusion question should be more demanding: What specific role does this element play in answering the research question or implementing the design?
A variable may be needed because the theoretical framework predicts its relationship with the outcome. A population may be necessary because the study compares groups. A setting may matter because the research problem is context-dependent. A data source may be essential because it provides evidence unavailable elsewhere.
If the only argument is "this is also related to the topic," the case for inclusion is usually weak.
Ask what changes if the element is removed
A practical way to evaluate a possible exclusion is to conduct a simple counterfactual test: imagine the study without it.
If you remove this population, variable, setting, comparison, period, method, or source of evidence, can you still answer the research question you actually wrote?
If yes, the element may be optional.
If no, either the element needs to remain or the research question itself needs to change.
This test is especially useful when a study has accumulated secondary questions and variables over time. Removing an element that contributes little to the central inquiry may improve focus. Removing a necessary comparison or explanatory variable may instead make the research question unanswerable.
Exclude populations that fall outside the population your question concerns
Population boundaries should follow from the target of inference or inquiry.
If the study investigates the transition into university, first-year students may be the appropriate population. Students in later years could have interesting experiences, but their inclusion is not automatically necessary.
If the study investigates differences between first-year and graduating students, however, excluding graduating students would eliminate one side of the comparison and undermine the research question.
Participant eligibility should therefore be tied to the characteristics needed to answer the question. In studies using formal inclusion and exclusion criteria, methodological guidance emphasizes that inclusion criteria should reflect key characteristics of the target population, while exclusion criteria should have substantive reasons rather than simply duplicating the inverse of inclusion criteria.
Population exclusions can also affect the range of people to whom findings may reasonably apply. The consequences of narrowing the population should therefore be considered alongside the rationale for doing so.
Do not confuse a delimitation with an exclusion criterion
The terms are related, but they operate at different levels.
Delimitation
A deliberate boundary around the overall study, such as focusing on first-year university students rather than students at every academic level.
Exclusion criterion
A rule used to determine that a potential participant, case, document, or other unit that might otherwise be considered will not be included because of a specified characteristic.
For example, defining the target population as first-year undergraduate students is a population delimitation. Within that population, a clinical study might exclude individuals with a condition that makes participation unsafe, while another design might exclude cases lacking the data necessary for the planned analysis.
Exact use of inclusion and exclusion criteria varies across methodologies, so the terminology should follow the research design rather than being applied mechanically to every study.
Exclude variables that do not have a clear conceptual or analytical role
Variables are particularly susceptible to accumulation because datasets and questionnaires can make adding them seem inexpensive.
But every additional construct creates obligations. Why is it measured? How is it defined? What theory or evidence connects it to the research problem? How will it enter the analysis? Does the design have adequate observations for the intended analysis? What will the resulting association mean?
Suppose your study investigates the relationship between generative AI use for academic writing and writing self-efficacy. You discover validated measures for technology acceptance, motivation, academic anxiety, critical thinking, creativity, digital literacy, satisfaction, and perceived usefulness.
The availability of instruments does not establish that all eight constructs belong in your study.
A useful variable should have a defensible role such as an exposure, outcome, predictor, confounder, mediator, moderator, control variable, or other theoretically meaningful construct appropriate to the design. If you cannot explain its role without saying "it might be interesting," consider leaving it out or treating it as a question for later research.
Do not exclude a variable merely because it complicates the expected result
There is an important difference between removing an unnecessary variable and excluding evidence that could challenge the conclusion you hope to reach.
Suppose previous research indicates that prior academic achievement is strongly associated with both the exposure and outcome in your proposed observational study. Excluding it simply because accounting for it makes the analysis more complicated could compromise the interpretation of the relationship you intend to estimate.
Likewise, removing inconvenient cases after seeing that they weaken an effect is not ordinary scope refinement. Depending on the circumstances, post hoc exclusion can introduce bias and undermine the credibility of the analysis.
Watch Out
A defensible exclusion narrows the question; it should not manufacture the answer. Do not remove populations, observations, variables, periods, or evidence merely because their inclusion makes the expected result weaker, messier, or less statistically convenient.
Exclude settings when they do not contribute to the question
More sites can increase contextual breadth, but each additional setting can also introduce variation that the study needs to understand.
A study evaluating implementation of one university's newly introduced AI policy may appropriately focus on that institution because the policy itself defines the case. Adding several universities with different policies would change the inquiry from an institutional implementation study into something closer to a comparative study.
Conversely, if the question asks whether implementation differs across institutional types, restricting the project to one institution would remove the comparison the question requires.
The decision should therefore follow the analytical purpose of the setting rather than an assumption that more locations automatically make research stronger.
Exclude time periods when they fall outside the phenomenon of interest
Temporal boundaries can also be methodologically meaningful.
If a study concerns institutional practice after a particular policy was introduced, records from years before the policy may not belong in the main analysis unless a pre-policy comparison is necessary. A study of student experiences during emergency remote teaching may need a period corresponding specifically to that educational condition.
On the other hand, excluding earlier years merely because they show a different pattern could distort a trend analysis.
The appropriate time boundary should therefore be determined by the phenomenon and design, not by which period produces the neatest findings.
Exclude outcomes that do not serve the study's central purpose
Researchers often want to measure every plausible benefit or consequence of an intervention. An educational technology study, for example, might consider achievement, engagement, motivation, satisfaction, self-efficacy, cognitive load, retention, attendance, creativity, and intention to continue using the tool.
Several outcomes can be justified when the study is designed to address them. The problem arises when outcomes accumulate without prioritization or adequate rationale.
In quantitative research, additional outcomes and comparisons can create further sample-size, statistical, and interpretation considerations. In any methodology, each outcome also expands the conceptual burden of the study.
Ask which outcome most directly represents the problem the study was designed to address. Secondary outcomes should have a reason for being secondary rather than simply joining the project because they are measurable.
Exclude methods that do not answer a necessary part of the question
Using more methods does not automatically make research more rigorous.
A researcher may feel that a survey should be supplemented by interviews because "mixed methods is stronger," or that observations should be added to an interview study because triangulation sounds desirable. Additional methods are valuable when they serve a methodological purpose, not merely when they increase the number of data sources.
If qualitative interviews are necessary to explain how participants interpret a quantitative finding, an integrated mixed-methods design may be appropriate. If the interviews answer an unrelated question, they may simply create another strand of research.
Methods should earn their place in the same way populations and variables do: by helping answer the research question.
Exclude evidence that does not match the evidential purpose of the study
Research based on documents, databases, publications, archives, or digital traces also requires deliberate boundaries.
A study of official university AI policies may include formally adopted policy documents while excluding informal social-media posts. A bibliometric study may specify document types, databases, subject areas, languages, or publication periods. A systematic or scoping review defines eligibility criteria governing which studies can contribute evidence.
In evidence synthesis, inclusion and exclusion criteria should be specified transparently because they determine the body of literature from which conclusions are drawn. Frameworks such as PICO, SPICE, or SPIDER may help structure eligibility depending on the review question and methodology, although no single framework fits every review.
The same principle applies beyond reviews: your source boundary should correspond to the evidence needed to answer the question.
Feasibility matters, but it should lead to a coherent question
Time, access, funding, expertise, equipment, data availability, and ethical constraints all affect what a study can realistically accomplish. Ignoring them can produce a research plan that looks impressive on paper but cannot be executed properly.
Feasibility can therefore justify narrowing.
However, the logic should not be: "I cannot obtain the evidence my question requires, so I will keep the same question and simply exclude that evidence."
Instead, revise the question and scope together until the evidence you can realistically obtain is capable of supporting the question you intend to answer.
If a nationwide study is impossible but a single-institution study is feasible, the solution may be a well-justified institutional study with appropriately bounded claims. It is not a single-institution dataset described as though it still represents the entire country.
Some exclusions improve depth rather than merely reducing workload
Leaving something out can improve a study when it allows greater attention to what remains.
A qualitative project may investigate one participant group deeply rather than several groups superficially. A quantitative study may test a clearly specified model rather than examining every variable in a large dataset. A case study may concentrate on one bounded implementation rather than making weak comparisons across unrelated cases.
This is why a deliberate exclusion should not automatically be treated as an apology. A well-chosen boundary can be part of methodological discipline.
The relevant question is whether narrowing the scope improves the study's coherence and evidential quality.
But narrowing can eventually remove too much
There is a point at which exclusion stops sharpening the study and begins hollowing it out.
Suppose a researcher narrows a project repeatedly: one institution, one course, one instructor, one assignment, one week, one outcome, and a very small group of participants. Such a study might still be valuable under an appropriate case-based or qualitative rationale. But if the original question seeks a broadly applicable estimate of an intervention's effectiveness, the accumulated restrictions may no longer support that purpose.
A narrow study is not automatically trivial, just as a broad study is not automatically important. The relevant issue is whether the remaining inquiry can still produce an answer with theoretical, empirical, methodological, practical, or contextual significance.
If each round of exclusion makes the project easier while making the answer progressively less consequential, it is worth asking whether the research question has become too trivial.
Consequential exclusions should be visible to the reader
Once you deliberately exclude something important, do not hide the decision.
Readers need enough information to understand the boundaries of the evidence. This is particularly important when the excluded population, variable, setting, period, or source could plausibly affect how the findings are interpreted.
A strong explanation usually does three things: identifies the boundary, provides the reason for it, and preserves the consequences of that boundary when interpreting the results.
For example: the study focuses on first-year students because the research problem concerns transition into university; students in later years are therefore outside the target population; conclusions should consequently not be written as though every undergraduate year level was investigated.
The more consequential the exclusion, the stronger the case for explaining why the population or variable was excluded.