03 · What You Need to Know
The Literature Gives You Candidates, Not a Ready-Made Variable List
Why researchers naturally look to previous studies
Looking at earlier studies is entirely sensible. The literature can reveal constructs that have repeatedly been associated with an outcome, variables used to address confounding, theoretically important mechanisms, common moderators, established baseline predictors, and measurement approaches that have already been tested.
Previous studies can therefore help answer questions such as:
- What variables have researchers considered important?
- What theoretical models have been used?
- What alternative explanations have been discussed?
- What factors may precede, mediate, or modify the focal relationship?
- What variables should be considered during study design?
The mistake occurs when this useful evidence is converted into a mechanical rule: they included it, so I should too.
Ask why the previous study included the variable
The same variable may appear in two studies for completely different reasons.
Suppose both papers include prior achievement.
One may use prior achievement to address confounding in an observational study of voluntary tutoring. Another may include it in a randomized trial because baseline achievement strongly predicts the outcome and improves precision.
The regression tables may look similar, but the rationale is different.
| Why a previous study included the variable |
What that means for your study |
|
Confounding control
|
Check whether the same causal structure applies to your exposure and outcome. |
|
Mediation
|
Ask whether the variable lies on the pathway you are studying. |
|
Moderation
|
Determine whether your theory predicts the same conditional relationship. |
|
Prediction
|
Assess whether the variable is available and useful for your prediction task. |
|
Precision
|
Consider whether the variable is prognostic and appropriate for your design. |
|
Descriptive reporting
|
Do not assume it belongs in the inferential model merely because it was measured. |
The same variable may have a different role in your study
Suppose digital literacy was treated as a confounder in an earlier observational study because it influenced both voluntary AI-tool use and academic performance.
Your study may instead evaluate a digital-literacy intervention and hypothesize that increased digital literacy leads to better learning outcomes.
Digital literacy is no longer playing the same role.
This illustrates why the same variable can be a confounder in one study and a mediator in another. Variable roles depend on the focal exposure, outcome, timing, and causal question.
Similar topic does not necessarily mean similar causal structure
Two studies may both examine “AI use and academic performance” while asking different questions.
One might compare voluntary users with nonusers.
Another might randomly assign access to an AI tutor.
A third might investigate which students choose to use the tool.
A fourth might predict final grades from AI-use behavior.
Those studies can involve the same constructs while requiring different variables.
For example, prior motivation may confound voluntary use, predict uptake, modify an intervention effect, or simply contribute to a prediction model. Its relevance changes with the question.
A statistically significant variable is not automatically a necessary variable
Researchers sometimes justify inclusion by stating that previous studies found the variable to be “significant.”
That is weak justification by itself.
A statistically significant association in another sample does not establish that the variable:
- is theoretically necessary in your model;
- is a confounder of your focal relationship;
- should be controlled;
- will improve prediction in your population;
- will replicate in your design;
- belongs in your conceptual framework.
Statistical significance is evidence about an estimate under a particular model and sample. It is not a permanent passport granting the variable entry into every subsequent study.
A nonsignificant variable may still matter
The reverse is also true.
Suppose a previous study included prior socioeconomic conditions as part of a theoretically justified confounding adjustment set, but its individual regression coefficient was nonsignificant.
That does not necessarily mean the variable was unnecessary.
Confounder selection should not be based simply on whether each covariate independently predicts the outcome at p <.05. Likewise, a theoretically central variable can produce an imprecise estimate in one study without becoming conceptually irrelevant.
The role of the variable matters more than whether one coefficient crossed a conventional significance threshold.
Repeated use may reflect convention rather than evidence
A variable can become standard within a literature because researchers repeatedly copy earlier models.
Eventually, everyone seems to control for age, sex, tenure, rank, firm size, socioeconomic status, or another familiar characteristic because everyone else does.
This creates a methodological inheritance problem: repetition can look like validation even when the original reason for inclusion has been forgotten.
Watch Out
Frequency of use in published studies does not prove that a variable is causally necessary. A common practice can be reasonable, outdated, context-specific, or simply reproduced by convention.
Look beyond the regression table
If you want to understand why a previous study included a variable, the regression table may not be enough.
Check:
- the research questions;
- the theoretical framework;
- the study design;
- the timing of measurement;
- the methods section;
- the stated covariate-selection strategy;
- the causal assumptions, if provided;
- the sensitivity or robustness analyses.
A variable that appears in Model 3 may have been included for sensitivity analysis rather than because the authors believed it belonged in the primary causal model.
Do not confuse adjustment variables with conceptual variables
Some variables appear in an analysis for technical or design reasons but are not central to the substantive theory.
For example, researchers may include study site, cohort, baseline outcome, stratification factors, or fixed effects. These variables may be analytically important without deserving a prominent box in the conceptual framework.
Conversely, a theoretically central mediator may appear prominently in the conceptual framework because it represents the mechanism being tested.
Copying every adjustment variable from a previous study into your conceptual framework therefore mixes two different purposes.
Previous inclusion should trigger a role question
When you encounter a recurring variable in the literature, ask:
What role was this variable supposed to play?
The possibilities include:
- focal predictor or exposure;
- outcome;
- antecedent;
- mediator;
- moderator;
- confounder;
- precision variable;
- predictor in a forecasting model;
- design variable;
- descriptive characteristic.
This role-based approach is more useful than simply categorizing variables as “used before” or “not used before.”
Confounders require causal justification
Suppose several previous papers controlled for age.
Should you?
If your study has a causal objective, ask whether age participates in a relevant noncausal pathway between your exposure and outcome.
If it does, age may belong in the adjustment set. If it does not, calling it a confounder simply because earlier authors adjusted for it is misleading.
This distinction follows from the fact that a control variable is not automatically a confounder.
Mediators should not be copied into “control” lists
Suppose an earlier study includes engagement because it investigates whether an intervention works through engagement.
If your objective is to estimate the total effect of that same intervention, mechanically controlling for engagement can remove part of the pathway through which the intervention operates.
Thus, a variable that was essential in the earlier mediation analysis may be inappropriate as a generic control in your model.
The difference becomes clear once you understand what it means for a mediating variable to explain part of a relationship.
Moderators should be included because you expect heterogeneity
If earlier researchers found that an intervention effect differed by age, should you test age as a moderator?
Possibly, but ask whether the moderation is theoretically plausible and relevant to your context.
An interaction discovered in one sample may not generalize. The earlier finding might also have been exploratory, imprecise, or dependent on how variables were coded.
A stronger rationale combines prior evidence with a substantive explanation for why the X–Y relationship should differ across values of the proposed moderator.
Replication can be a valid reason to retain the same variable set
There are cases in which reproducing an earlier variable specification is exactly the point.
If you are conducting a direct or close replication, keeping the earlier model may be necessary to evaluate whether the original result reproduces under comparable conditions.
Likewise, a planned comparative study may intentionally use the same covariate definitions across populations or time periods.
In those situations, prior inclusion is not being followed blindly. It is part of the research design.
Copying variables by habit
“They used these variables, so we used them too.”
Replicating a specification deliberately
“We reproduced the original model because comparability with the prior study is itself part of the research question.”
Comparability across studies can also justify consistent variables
Researchers conducting multi-site studies, longitudinal comparisons, meta-analytic harmonization, benchmarking, or repeated institutional analyses may need consistent variable definitions.
Consistency can improve comparability, but it still helps to distinguish this purpose from causal necessity.
A variable may be retained because comparability matters, not because it is theoretically indispensable in every individual analysis.
Measures may not transfer even when constructs do
Suppose previous research measured “technology readiness” among corporate employees. You are studying first-year university students.
The construct may remain relevant, but the original instrument, operationalization, or cut points may not transfer appropriately.
Before importing a variable from the literature, ask:
- Is the construct relevant to this population?
- Does the measure have suitable validity evidence?
- Does the wording fit the context?
- Is the timing appropriate?
- Does the construct mean the same thing across groups?
Variable selection and measurement selection are related but separate decisions.
Context can change the meaning of a variable
Institutional support in a well-resourced private university may not operate in the same way as institutional support in a resource-constrained public institution.
Likewise, socioeconomic status, technology access, academic rank, or workload can have different distributions and consequences across settings.
The fact that a variable predicted an outcome elsewhere does not guarantee the same relationship in your population.
This does not mean ignoring previous research. It means treating generalizability as a question rather than an assumption.
Theoretical saturation is not the goal of a single study
Researchers sometimes worry that omitting any variable mentioned in prior research makes their model incomplete.
But no single study needs to represent every plausible determinant of an outcome.
A conceptual framework should answer a bounded research question. Variables outside that boundary can remain important without belonging in the present analysis.
This is part of the broader challenge of deciding which variables actually belong in your study.
Ask whether inclusion changes the study's scientific question
Adding a variable can do more than make the model “more controlled.” It can change what is being estimated.
If a mediator is added, the coefficient for X may shift from something closer to a total relationship toward a conditional direct relationship.
If a moderator and interaction are added, the study moves from one average X–Y relationship to conditional relationships.
If an antecedent is added, the framework begins explaining where X itself comes from.
Every new variable can therefore expand or redirect the research question.
Use previous studies to build a candidate-variable table
A practical literature-review strategy is to record not only which variables appear, but why.
| Candidate variable |
Role in previous study |
Relevant to my question? |
Proposed role in my study |
| Prior achievement |
Confounder |
Yes |
Potential confounder |
| Age |
Descriptive/control |
Unclear |
Needs justification |
| Engagement |
Mediator |
Yes |
Proposed mediator |
| Institutional support |
Moderator |
Outside primary question |
Exclude from primary model |
This forces the literature review to inform reasoning rather than merely accumulate citations.
The absence of a variable from previous studies does not make it illegitimate
Research also progresses by identifying relationships earlier studies did not examine.
If strong theory, qualitative evidence, new technology, changed policy, or a different population suggests that a previously neglected variable matters, its novelty can be scientifically valuable.
Novel inclusion should still be justified. But “no one has used this variable before” is not, by itself, a reason to exclude it.
Previous research is a foundation, not a boundary fence.