Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Decide Which Variables Actually Belong in Your Study?

Variables should earn their place in a study by helping answer the research question, represent the theory, address the design, or support the intended analysis. More variables do not automatically produce a stronger study.

109
Choosing Variables for a Research Study Guide 109 of 223
01 · The Question

Out of All the Variables You Could Study, Which Ones Actually Belong?

Once researchers begin reading the literature, the number of potentially relevant variables can grow quickly. One paper identifies motivation. Another includes age and socioeconomic status. A third adds self-efficacy, prior achievement, institutional support, engagement, digital literacy, personality, and several demographic characteristics.

Soon the conceptual framework starts to resemble a railway map.

Which of these variables should you actually include?

The answer is not “everything that has been significant before,” “everything available in the dataset,” or “as many variables as possible.” Each variable should have a defensible role in relation to the research question, theoretical framework, causal structure, design, measurement strategy, or analytical objective.

Choosing variables is therefore a substantive research decision before it becomes a statistical one. A study becomes stronger when the variables form a coherent answer to a focused question, not when the model simply contains more boxes and arrows.

02 · The Short Answer

Every Variable Should Have a Clear Reason for Being in the Study

In Brief

Choose variables by starting with the research question and then including only the constructs needed to represent the focal relationship, theoretical explanation, causal structure, design requirements, or analytical purpose.

A variable should not be included merely because it is available, statistically significant in earlier studies, or commonly used as a control. For each candidate variable, you should be able to explain what role it plays, why that role matters, how it will be measured, and what decision or inference would change if the variable were omitted.

03 · What You Need to Know

Variable Selection Begins With the Question, Not the Dataset

Start with the outcome or phenomenon you are trying to explain

A useful first step is to identify the phenomenon that makes the study worth conducting.

What are you trying to understand, compare, predict, explain, or change?

That may be:

  • academic achievement;
  • technology adoption;
  • employee retention;
  • student persistence;
  • research productivity;
  • health behavior;
  • organizational performance.

Once the focal outcome is clear, ask what relationship or process the study is actually investigating.

For example:

Does AI teaching self-efficacy predict faculty adoption of generative AI?

This already identifies two central variables: self-efficacy and adoption.

The remaining variables should be added only if they help answer a more specific theoretical, causal, predictive, or design question.

Define the focal relationship before adding third variables

Researchers sometimes begin with a list of constructs and then try to connect them.

A stronger strategy is to begin with a focal relationship:

X → Y

or, if the study is associational:

X ↔ Y

Then ask what is missing from the explanation.

Do you need to understand:

  • what causes X?
  • how X affects Y?
  • when the X–Y relationship changes?
  • whether another variable distorts the X–Y comparison?
  • whether additional variables improve prediction?

Each of those questions justifies a different type of variable.

Every variable should have a role

Possible role Question it helps answer Example
Exposure / focal predictor What variable is central to the relationship being studied? AI teaching self-efficacy
Outcome What phenomenon are you trying to explain or predict? AI classroom adoption
Antecedent What contributes to the focal predictor? Professional development
Mediator Through what process does X affect Y? Perceived usefulness
Moderator When or for whom does the X–Y relationship change? Institutional support
Confounder What may bias the causal comparison between X and Y? Prior digital competence
Predictor / feature What information improves forecasting? Multiple behavioral indicators

The labels are not permanent properties of the variables. Their meaning depends on the research question, which is why the same variable can occupy different causal roles in different studies.

Ask whether the study is associational, causal, explanatory, or predictive

The purpose of the study strongly influences variable selection.

If the goal is association, you may focus on describing relationships among variables without claiming that changing one causes changes in another.

If the goal is causal inference, variable selection must reflect the causal structure and the effect being estimated.

If the goal is mediation or mechanism, you need variables representing theoretically plausible pathways.

If the goal is moderation, you need variables that plausibly define heterogeneity in the focal relationship.

If the goal is prediction, variables may be valuable because they improve out-of-sample forecasting even when they are not causal.

This follows from the distinction between association, effect, and prediction. A variable useful for one purpose may be unnecessary or even problematic for another.

Theoretical importance is stronger than statistical convenience

A variable belongs in a theoretical model because there is a defensible reason to believe it plays the proposed role.

That reason may come from:

  • an established theory;
  • a well-supported conceptual model;
  • prior empirical evidence;
  • known temporal ordering;
  • substantive domain knowledge;
  • a clearly articulated causal hypothesis.

Statistical significance in another paper can contribute evidence, but it should not be the sole reason for inclusion.

Likewise, a variable does not become theoretically important simply because your dataset happens to contain it.

Do not confuse literature frequency with conceptual necessity

If age appears in twenty studies, that tells you researchers often considered age relevant. It does not establish that age belongs in your analysis.

Those studies may have:

  • asked different questions;
  • used age for descriptive purposes;
  • included it for precision;
  • treated it as a moderator;
  • controlled for it by convention;
  • used a different outcome;
  • studied another population.

The literature helps you identify candidate variables. Your research question determines which candidates survive.

This is why including a variable because previous studies did requires more justification than citation count.

Use theory to decide direction, not just inclusion

Choosing a variable is only half the problem. You also need to decide how it relates to the others.

Suppose self-efficacy and engagement are associated. Should the framework show:

Self-efficacy → Engagement

or:

Engagement → Self-efficacy

or perhaps a reciprocal relationship?

Direction should be informed by theory, temporal evidence, prior longitudinal findings, experimental evidence where available, and substantive plausibility.

If the evidence genuinely does not establish direction, researchers should consider what to do when relationship direction is unclear rather than drawing an arrow simply because the software requires one.

A conceptual framework should not be a catalogue of variables

A conceptual framework should represent the relationships central to the study.

Adding a box for every construct mentioned in the literature creates complexity without necessarily creating explanation.

For each candidate variable, ask:

What proposition does adding this variable allow the study to test that it could not test without it?

If the answer is unclear, the variable may not belong.

This criterion is particularly useful when determining whether a proposed relationship belongs in the conceptual framework.

Some variables belong in the analysis but not in the conceptual framework

Not every variable included in a statistical model needs a prominent box in the conceptual framework.

A study may adjust for site, cohort, baseline outcome, age, or sampling factors for analytical reasons while keeping the conceptual framework focused on the substantive theory.

Conversely, a conceptual framework may include theoretical constructs whose operationalization requires multiple indicators rather than one simple measured variable.

Conceptual variables Represent the substantive constructs and relationships central to the theory or research question.
Analytical variables May also include adjustment variables, design factors, technical covariates, or other variables needed for estimation.

The two sets can overlap substantially without being identical.

For causal studies, think in terms of causal structure

If the study intends to estimate a causal effect, variable selection should not rely only on regression conventions.

Ask:

  • What causes the exposure?
  • What causes the outcome?
  • Which variables are common causes of both?
  • Which variables occur after the exposure?
  • Which variables lie on the causal pathway?
  • Which variables could create selection problems if conditioned on?

A directed acyclic graph, or DAG, can help represent these assumptions.

The purpose is not to draw an impressive diagram. It is to identify which paths should be blocked, which should remain open, and what set of variables is sufficient for the causal effect being estimated.

Do not treat every plausible confounder as a generic control

If a variable is introduced to address confounding, its justification should be causal.

A control variable is not automatically a confounder, and confounders should not be selected merely by looking for significant correlations with the exposure and outcome.

The inclusion decision should follow from the assumed causal structure.

Do not automatically control for mediators

If X affects M and M affects Y, M lies on the pathway:

X → M → Y

If your target is the total effect of X on Y, adjusting for M may remove part of the effect you want to estimate.

If your question concerns mediation, M should be modeled as a mediator under a suitable framework.

This illustrates a broader rule: the same statistical operation, adding a variable to a model, can be appropriate or inappropriate depending on why the variable is there.

Moderators require a reason to expect heterogeneity

A moderator should not be selected simply by testing dozens of interaction terms and retaining whichever one produces p <.05.

A stronger moderation hypothesis explains why the focal relationship should differ by the moderator.

For example, an intervention may be expected to work better among beginners because advanced learners have less room to improve. The moderator has a theoretical rationale.

Without such reasoning, interaction searches can become exploratory fishing exercises disguised as conceptual sophistication.

Antecedents are useful when the study asks where a focal variable comes from

Suppose self-efficacy is the focal predictor of adoption. If the research question also asks what shapes self-efficacy, variables such as prior mastery experiences or professional development may be justified as antecedents.

But adding every possible upstream determinant can expand the study indefinitely.

An antecedent variable belongs when explaining its downstream construct is part of the actual research objective.

Prediction follows a different variable-selection logic

In prediction, a variable may be included because it improves forecasting rather than because it is causally upstream.

A hospital readmission model, dropout-risk model, or academic-performance predictor may legitimately use variables that function as proxies or markers rather than causes.

The key questions become:

  • Does the variable improve out-of-sample performance?
  • Is it available at the time prediction must be made?
  • Is the measurement stable and reliable?
  • Does using it create fairness, leakage, or implementation concerns?

A prediction model therefore should not be judged by the same covariate-selection logic as a causal model.

Measurement quality can eliminate a theoretically attractive variable

A variable can be theoretically important but practically unusable if it cannot be measured adequately.

Suppose theory identifies instructional quality as a central moderator, but the study has only a single poorly defined self-report item. Including the variable may add apparent complexity without producing a credible test of the theoretical proposition.

Before including a variable, ask:

  • Can the construct be measured validly?
  • Is the measure reliable enough for the intended inference?
  • Does the measure capture the construct at the correct time?
  • Is the measure comparable across groups if group comparisons are intended?

Conceptual importance and measurement quality must meet somewhere in the middle.

Timing can determine whether a variable belongs at all

A variable measured after the outcome cannot ordinarily explain that earlier outcome as a causal antecedent.

A mediator should occur in the relevant sequence after exposure and before outcome.

A predictor in a prospective prediction model must be available before the prediction is needed. Otherwise, the model may suffer from information leakage.

Timing is therefore not a technical afterthought. It helps determine whether the variable can perform the role assigned to it.

Feasibility matters, but it should not silently redefine the question

Researchers often face practical constraints: participant burden, instrument length, access to administrative records, sample size, cost, ethical restrictions, or time.

Some theoretically relevant variables may therefore be impossible to measure.

The appropriate response is not to pretend they do not matter. Instead, researchers should:

  • narrow the research question when necessary;
  • acknowledge important unmeasured variables;
  • avoid causal claims that require unavailable adjustment;
  • consider alternative designs or proxies cautiously.

Feasibility shapes the study, but transparency should accompany the compromise.

Sample size limits how complicated the model can reasonably become

Every added variable introduces parameters to estimate. Interactions, nonlinear terms, latent constructs, mediation pathways, and repeated measures can increase the demand further.

A model containing many weakly justified variables may be unstable, overfit the sample, inflate standard errors, or make interpretation difficult.

There is no universal rule stating that a particular number of observations allows exactly a particular number of variables across all statistical methods. Sample-size planning should reflect the design, expected effect sizes, parameterization, measurement model, missingness, and inferential objective.

The larger lesson is simpler: model complexity should be planned rather than discovered accidentally after data collection.

More variables can make the study weaker

Adding variables may feel like increasing rigor, but unnecessary complexity can introduce several problems:

  • unclear hypotheses;
  • multiple-testing burden;
  • unstable estimates;
  • multicollinearity;
  • measurement burden;
  • participant fatigue;
  • missing data;
  • overadjustment;
  • difficulty explaining what the model actually tests.

This is why adding more variables can make a study weaker.

Do not choose variables by stepwise significance testing when theory defines their role

Automated procedures that add or remove variables based on p-values can sometimes have uses for particular predictive objectives, but they are poorly suited to deciding causal or theoretical roles.

A confounder does not stop being part of the assumed causal structure because its sample coefficient is nonsignificant. A mediator does not become theoretically irrelevant because one path misses a threshold. A moderator should not be invented because an exploratory interaction happened to be significant.

Statistical output should evaluate a research model, not retroactively write its theory.

Ask what would happen if the variable were removed

A powerful practical test is to consider what scientific claim would become impossible if the candidate variable were omitted.

If removing M prevents you from testing the proposed mechanism, M probably has a clear role.

If removing W means you can no longer test the hypothesized conditional effect, W has a reason to remain.

If removing Z does not change the research question, identification strategy, predictive objective, design, or interpretation, perhaps Z does not need to be there.

This test is not infallible, but it forces the variable to justify its place.

Separate “must include,” “useful to include,” and “interesting but outside scope”

Not all candidate variables deserve the same priority.

Category Meaning Typical action
Must include Essential to the focal question, identification strategy, design, or primary theory Include and justify clearly
Useful to include Adds precision, a prespecified secondary question, or meaningful robustness information Include when feasible and analytically justified
Interesting but outside scope Potentially relevant but not necessary to answer the present question Leave for future research rather than expanding the model indefinitely

This distinction can prevent scope creep while preserving awareness of the wider literature.

The strongest conceptual frameworks are selective

A good conceptual framework does not prove sophistication by density.

It should make the central theoretical propositions visible.

If the framework requires several minutes of tracing arrows before anyone can identify the main question, the problem may not be that the audience lacks methodological sophistication. The model may simply be trying to do too much.

Selectivity is therefore part of rigor. Removing a weakly justified variable can improve theoretical clarity, measurement quality, analysis, and interpretation all at once.

04 · A Practical Example

How to Reduce a Long List of Variables to a Coherent Study

Hypothetical Example

Faculty adoption of generative AI for teaching

A researcher wants to understand why university faculty adopt generative AI in teaching. The literature identifies digital competence, AI self-efficacy, perceived usefulness, perceived ease of use, institutional support, age, academic rank, teaching experience, disciplinary field, professional development, technology anxiety, innovativeness, workload, ethical concerns, and peer influence.

Define the focal question The researcher decides to study whether AI teaching self-efficacy is associated with adoption and whether perceived usefulness helps explain that relationship.
Identify essential variables AI teaching self-efficacy becomes X, adoption becomes Y, and perceived usefulness becomes the proposed mediator M.
Identify relevant adjustment variables The researcher considers whether prior digital competence may plausibly contribute to both AI self-efficacy and adoption. If the causal assumptions support that structure, digital competence may require consideration as part of the adjustment strategy.
Evaluate additional constructs Institutional support is theoretically interesting, but if the study is not designed to investigate effect heterogeneity or institutional context, it may be outside the primary model rather than included merely because prior studies measured it.
Set boundaries Age, rank, discipline, workload, innovativeness, anxiety, ethical concern, and peer influence are not automatically added. Each is retained only if it serves a clear design, causal, theoretical, or secondary analytical purpose.

The result is a more focused study, not a less sophisticated one.

The researcher can explain why every major variable appears in the model and what question each variable helps answer. Variables excluded from the primary model can still be acknowledged as plausible influences beyond the study's scope.

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing Research Variables

Misconception

If a variable was significant in previous studies, it belongs in mine

Not necessarily. Previous statistical significance does not establish that the variable is relevant to your research question, causal structure, population, or design. Prior evidence should inform variable selection rather than dictate it mechanically.

Misconception

More variables make a conceptual framework more comprehensive

More variables can make a framework less coherent. A comprehensive answer to a focused question is different from a model that attempts to represent every factor mentioned in the literature.

Misconception

Every demographic variable should be included as a control

No. Demographic variables should have a defensible analytical, causal, descriptive, or design purpose. Age, sex, rank, or socioeconomic status do not become necessary controls simply because they are easy to collect.

Misconception

You can decide which variables matter after running correlations

Correlations can inform exploratory analysis but cannot determine causal roles or theoretical relevance. Variable selection should normally be guided substantially by the research question, theory, design, timing, and intended inference before the final model is fitted.

Misconception

If a variable is theoretically interesting, it must be included

No. Many interesting constructs may lie outside the scope of a particular study. The relevant question is whether the variable is needed to answer the study's focal question and whether it can be measured and analyzed credibly.

Misconception

If the dataset already contains a variable, there is little harm in adding it

Availability is not justification. Additional variables can complicate interpretation, increase missing-data losses, introduce inappropriate adjustment, inflate multiplicity, and encourage post hoc storytelling.

Misconception

A complex model is automatically more publishable

Complexity may be warranted when the research question is genuinely complex, but reviewers can also recognize decorative complexity. A smaller model with clear theoretical justification and appropriate measurement is often more defensible than a crowded framework whose paths are difficult to explain.

06 · What This Means for You

Make Every Variable Earn Its Place

A practical way to choose variables is to require a short justification for each one before data analysis begins.

A simple decision framework

If the variable defines the focal exposure, predictor, or outcome
Include it because it is essential to the primary research question.
If the variable represents a theoretically proposed mechanism or boundary condition
Include it when mediation or moderation is genuinely part of the research question.
If the variable is needed for causal identification, design adjustment, or prespecified precision improvement
Include it for that explicit analytical reason and report the rationale.
If the only justification is that the variable exists in the dataset or appeared in several papers
Do not include it automatically. Determine what role, if any, it plays in your study.
If the variable is theoretically interesting but peripheral to the main question
Consider excluding it from the primary model and treating it as a future or secondary research question.

For each retained variable, you should ideally be able to complete this sentence:

“This variable is included because…”

If the explanation is theoretically, causally, methodologically, or predictively meaningful, the variable has a defensible role. If the sentence ends with “because previous papers had it” or “because it was in the dataset,” the rationale probably needs more work.

07 · A Quick Checklist

Before Finalizing Your Variables, Check This

For every candidate variable, check:
Can I state exactly what research question this variable helps answer?
Can I identify its role as exposure, outcome, antecedent, mediator, moderator, confounder, predictor, design variable, or another justified category?
Is the proposed relationship supported by theory, prior evidence, or defensible substantive reasoning?
Is the temporal ordering compatible with the role assigned to the variable?
Can the construct be measured with adequate validity and reliability?
Does including the variable change the causal estimand or create inappropriate adjustment?
Is the model complexity realistic for the design and available sample?
Would removing the variable prevent the study from answering an important part of its stated question?
Am I including the variable for a substantive reason rather than because it was available or previously significant?
Can I explain the final model clearly without turning the conceptual framework into an inventory of the literature?
08 · Frequently Asked Questions

Frequently Asked Questions About Choosing Research Variables

How many variables should a research study have?

There is no universal ideal number. The appropriate number depends on the research question, theoretical model, study design, measurement burden, sample size, analytical method, and intended inference. Include enough variables to answer the question credibly, but not additional variables merely to make the model appear comprehensive.

Should I include every variable identified in the literature review?

No. The literature review identifies potentially relevant constructs, but the study should retain only those needed for its specific theoretical, causal, predictive, or design purpose.

Should demographic variables always be controls?

No. Demographic variables should be included when they have a defensible role in the causal model, descriptive objective, design, precision strategy, subgroup analysis, or prediction task. They should not be treated as automatic controls.

Can I choose variables based on correlations?

Correlations can be useful descriptively and exploratorily, but they should not by themselves determine causal roles or theoretical inclusion. Strongly correlated variables may be irrelevant to the focal question, while theoretically essential variables may show modest sample correlations.

How do I know whether a variable should be a mediator or moderator?

Ask what theoretical question you are posing. A mediator lies on a proposed pathway through which X affects Y. A moderator indicates that the X–Y relationship differs depending on another variable. The distinction should be specified before choosing the corresponding statistical analysis.

Should I include a variable because a reviewer might expect it?

Anticipating reasonable methodological concerns is useful, but a variable still needs a defensible role. It is better to explain why an unnecessary variable was not included than to add it without understanding how adjustment affects the analysis.

What if an important variable cannot be measured?

Consider whether the research question needs to be narrowed, whether a defensible proxy exists, whether another design can address the issue, and how the missing variable limits interpretation. Do not silently make causal claims that depend on information the study cannot observe.

Can I add variables after seeing the results?

Exploratory analyses can be scientifically useful, but they should be identified as exploratory. Adding variables solely because they improve significance or produce a more appealing story risks overfitting and post hoc interpretation. Prespecifying important roles strengthens confirmatory inference.

09 · The Bottom Line

A Good Study Includes the Variables Needed to Answer Its Question, Not Every Variable That Could Matter

The Bottom Line

Variables belong in a study when they have a clear role in answering the research question, representing the theory, identifying the causal effect, supporting the design, or improving the intended prediction task.

Do not equate model size with rigor. Start with the focal question, assign each variable a defensible role, check timing and measurement, consider analytical consequences, and remove variables that add complexity without adding a necessary scientific proposition.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes