03 · What You Need to Know
Every Additional Variable Has a Scientific and Statistical Cost
More variables can make the research question less clear
Consider a study that begins with a focused question:
Does academic self-efficacy predict students' persistence?
Then the researcher adds motivation, engagement, social support, socioeconomic status, digital competence, institutional climate, academic anxiety, prior achievement, learning strategies, instructor support, age, sex, academic program, and several interactions.
At some point, the study may no longer have one identifiable question.
A conceptual model becomes difficult to defend when variables accumulate without a corresponding theoretical argument explaining why each belongs.
Comprehensiveness and focus are not opposites. A focused study can address its question comprehensively without attempting to model every plausible cause of the outcome.
Complexity should follow the research question
A complicated question may genuinely require a complicated model.
If theory predicts that X affects Y through two mediators and that one pathway differs according to W, a complex model may be necessary.
But complexity should emerge from the substantive argument:
Question → theory → variables → model
rather than:
Available variables → model → post hoc story
This is why deciding which variables actually belong in the study should occur before fitting increasingly elaborate models.
Every variable creates measurement burden
A variable is not just another name in a regression equation. It has to be operationalized and measured.
If the study uses questionnaires, adding constructs can mean:
- more items;
- longer completion time;
- participant fatigue;
- straight-lining or careless responses;
- greater attrition;
- more missing data.
Adding ten constructs with six items each means sixty additional responses per participant. A theoretically ambitious survey can quietly become a test of participant endurance.
Measurement burden therefore belongs in variable-selection decisions from the beginning.
Weak measurement does not become strong because there are many variables
Researchers sometimes respond to model complexity by measuring each construct more superficially.
Instead of using a defensible multi-item measure, they include one or two convenient items because the questionnaire is already too long.
The result may be a model containing many constructs measured poorly.
That trade-off can undermine the theoretical sophistication the larger model was supposed to create.
Fewer variables measured well may provide a stronger test than many variables represented weakly.
More predictors consume statistical information
Every estimated coefficient requires information from the data.
As the number of parameters grows relative to the amount of information in the sample, estimates can become less stable and more uncertain.
The issue is not simply the raw number of variables. A categorical predictor with several levels requires multiple parameters. Interactions introduce additional parameters. Nonlinear terms add more. Latent-variable models may estimate loadings, variances, covariances, structural paths, and measurement errors.
Thus, “ten variables” can represent very different levels of statistical complexity depending on how they enter the model.
There is no universal observations-per-variable rule
Rules such as “10 observations per predictor” are often repeated as though they apply universally.
They do not.
Sample requirements depend on the statistical model, outcome distribution, number of parameters, predictor distributions, expected effect sizes, event frequency for binary outcomes, missingness, penalization, validation strategy, and inferential goal.
A simple linear model with well-behaved continuous variables differs from logistic regression with rare outcomes, survival analysis, multilevel modeling, structural equation modeling, or machine-learning prediction.
Rather than applying one universal ratio, sample-size and model-complexity planning should reflect the actual analysis.
Too many predictors can increase overfitting
Overfitting occurs when a model captures sample-specific noise in addition to genuine signal.
An overfitted model may describe the development dataset very well but perform poorly on new observations.
This is especially important in prediction. Adding candidate predictors can improve apparent in-sample fit almost automatically, while out-of-sample performance may fail to improve and can deteriorate.
In-sample fit
How well the model describes the data used to build it.
Generalization
How well the model performs on new data that were not used to fit it.
Prediction models therefore require appropriate internal or external validation rather than judging added variables only by how much they improve fit in the original sample.
More variables can reduce precision
Adding covariates can sometimes improve precision, especially when they explain substantial outcome variation.
But additional predictors can also increase uncertainty, particularly when they are strongly correlated with one another or contribute little independent information.
Suppose several measures capture closely related aspects of digital competence. Their individual coefficients may become unstable because the model is trying to separate highly overlapping information.
Standard errors may increase, coefficient signs may change across specifications, and interpretations can become highly model-dependent.
Multicollinearity complicates interpretation
Multicollinearity occurs when predictors contain substantial overlapping information.
It does not necessarily bias ordinary least-squares coefficient estimates by itself, but it can make individual coefficients difficult to estimate precisely and interpret reliably.
If self-efficacy, confidence, perceived competence, and technology readiness are all highly related, including all four may make it difficult to claim that one coefficient represents a clearly distinct construct.
The problem may be statistical, conceptual, or both.
Before adding several closely related variables, ask whether they genuinely represent distinguishable theoretical constructs.
Adding variables creates more opportunities for chance findings
Suppose a study tests twenty predictors, fifteen subgroup effects, ten interactions, and several alternative outcomes.
The number of possible findings becomes large.
If researchers selectively highlight only statistically significant results, the probability of presenting chance patterns as discoveries increases.
This issue is not solved merely by reporting p-values correctly. Researchers should distinguish confirmatory hypotheses from exploratory analyses, consider multiplicity where relevant, and avoid designing a conceptual framework retrospectively around whichever paths happened to be significant.
Every moderator multiplies the number of relationships you must interpret
Adding a moderator is not equivalent to adding one ordinary predictor.
A moderation hypothesis introduces an interaction and makes the focal effect conditional.
If several moderators are added, the number of possible conditional relationships can become difficult to communicate and may demand substantially larger samples to estimate precisely.
This is why a moderating variable should represent a meaningful conditional hypothesis, not simply another opportunity to search for significant interactions.
Every mediator adds causal assumptions
Mediation models can also become complex quickly.
Suppose researchers propose four parallel mediators between X and Y.
Each mediator requires justification for:
- its temporal position;
- the X → M pathway;
- the M → Y pathway;
- potential confounding;
- measurement quality;
- relationships among the mediators.
Adding mediators therefore does more than make a path diagram larger. It adds assumptions about how the process works.
A smaller number of theoretically compelling mediators can provide a clearer mechanistic test than a large set selected because all were measured.
More “controls” can create overadjustment
One of the most consequential misconceptions is that causal estimates improve as researchers control for more variables.
Suppose:
X → M → Y
If M is a mediator and the target is the total effect of X on Y, adjusting for M blocks part of the causal effect.
Other variables can create additional problems if they are colliders or descendants of variables affected by the exposure.
Watch Out
More adjustment is not automatically better adjustment. In causal research, conditioning on the wrong variables can change the estimand or introduce bias rather than remove it.
This is why a control variable should not automatically be treated as a confounder.
Adding post-exposure variables can answer a different question
Suppose researchers want the total effect of an educational intervention on achievement but adjust for motivation measured after the intervention.
If the intervention changes motivation, and motivation affects achievement, the adjusted model no longer estimates the same total-effect relationship.
Researchers can easily interpret this as “a more fully controlled estimate” when it is actually a different estimand.
Model expansion can therefore change the scientific question without making that change obvious.
More variables can increase missing-data problems
Suppose complete-case analysis requires every participant to have observed values on all variables in the model.
Adding more variables creates more opportunities for at least one value to be missing.
A sample of 1,000 participants can shrink substantially once the analysis requires complete information across twenty measures.
This can reduce precision and, depending on why data are missing, introduce selection bias.
Missing-data methods may mitigate some problems, but unnecessary variables can still make the missingness structure more complicated.
More constructs can increase participant attrition
Long questionnaires are not merely inconvenient. Participants may stop responding, skip matrix items, provide less thoughtful answers, or decline future waves of a longitudinal study.
This is particularly consequential when the variables generating extra burden are peripheral to the primary question.
Variable selection therefore has an ethical dimension. Researchers should not ask participants to provide information that is unlikely to contribute meaningfully to the stated scientific objectives.
Model complexity can undermine interpretability
A model may fit well statistically while becoming difficult to explain scientifically.
Imagine a structural model containing six antecedents, three mediators, four moderators, two outcomes, eleven covariates, and multiple correlated residuals.
Even if software successfully estimates the model, the researcher still has to explain what has been learned.
Which relationships were primary?
Which were prespecified?
Which findings change theory or practice?
If the answer requires navigating dozens of coefficients, the model may be analytically richer than the research question requires.
Complex conceptual frameworks can hide weak theory
A dense framework can create an appearance of theoretical sophistication. Yet every arrow should represent an actual proposition.
If the researcher cannot explain why X should affect M, why M should affect Y, and why W should change that relationship, the problem is not solved by adding citations around the diagram.
A useful test is to ask whether each relationship genuinely belongs in the conceptual framework.
Additional variables can make replication harder
The more decisions a study requires, the more difficult exact reproduction becomes.
Researchers need to know:
- which variables were included;
- how each was coded;
- which transformations were applied;
- which interactions were tested;
- which covariates were retained;
- which models were considered;
- which results were treated as primary.
A focused, prespecified model is often easier to reproduce and scrutinize than a large analytical garden containing many plausible paths.
Complexity can increase researcher degrees of freedom
With many variables, researchers gain many analytical choices.
Should age be continuous or categorical?
Should motivation be retained?
Which interaction should be tested?
Should the mediator be entered before or after another covariate?
Should one outlier be removed?
Should one subgroup be analyzed separately?
Each choice may be individually defensible, but a large number of post hoc choices can make results sensitive to researcher decisions.
Prespecification, transparent reporting, robustness analyses, and clear separation of exploratory from confirmatory work become more important as model flexibility increases.
More variables may shift attention away from the primary effect
A study may begin with a meaningful substantive question and end with discussion dominated by secondary covariates.
For example, the main purpose may be to evaluate an intervention, yet the discussion becomes occupied with unexpected age, rank, and discipline coefficients because they happen to be significant.
Secondary variables can provide useful context, but they should not automatically displace the question the study was designed to answer.
Not every variable that improves R² improves the study
In ordinary regression, adding predictors cannot decrease the unadjusted R² in the same sample.
This can create the misleading impression that every added variable improves the model.
But in-sample variance explained is only one criterion.
The new variable may:
- have no theoretical relevance;
- provide negligible predictive gain on new data;
- increase measurement burden;
- reduce interpretability;
- introduce causal overadjustment;
- require information unavailable in practice.
A higher in-sample R² therefore does not automatically mean a stronger study.
Adjusted R² and information criteria address only part of the problem
Statistical criteria such as adjusted R², AIC, BIC, cross-validation, or penalized regression can help evaluate model complexity for particular purposes.
They cannot determine whether the variables are theoretically meaningful or causally appropriate.
A model-selection criterion can prefer a statistically efficient model while remaining silent about whether a variable was measured after the outcome, represents a collider, or lacks construct validity.
Statistical model selection and scientific variable selection overlap but are not interchangeable.
Prediction may justify many variables, but validation becomes essential
Prediction models often consider many candidate predictors because the objective is accurate forecasting rather than estimating a small set of interpretable causal effects.
Even there, unrestricted variable addition is not automatically beneficial.
Researchers may use penalization, shrinkage, dimensionality reduction, or other strategies to manage complexity. Most importantly, model performance should be evaluated on observations not used to estimate the apparent fit.
The logic differs from causal research, reinforcing why prediction and causal explanation should not be conflated.
Some additional variables genuinely strengthen a study
The lesson is not “fewer variables are always better.”
An additional variable can substantially improve a study when it:
- addresses an important confounding path;
- represents a theoretically essential mediator;
- tests a meaningful boundary condition;
- improves predictive performance;
- accounts for a key design feature;
- improves precision;
- allows an important competing explanation to be tested.
The criterion is contribution, not count.
Parsimonious does not mean simplistic
Parsimony is sometimes interpreted as using the fewest variables possible.
A more useful interpretation is using no more complexity than the research question requires.
A parsimonious model can still include mediation, moderation, nonlinear relationships, repeated measurements, or multiple outcomes if those elements are necessary to represent the phenomenon adequately.
Removing essential complexity merely to produce a small model is no more defensible than adding unnecessary complexity to make a study look sophisticated.
Classify variables by priority before data collection
A useful planning strategy is to separate variables into categories.
| Priority |
Description |
Typical treatment |
|
Essential
|
Required for the primary research question, theory, design, or causal identification |
Include and measure carefully |
|
Secondary
|
Supports prespecified additional questions or precision |
Include when justified and feasible |
|
Exploratory
|
Potentially informative but not central to confirmatory claims |
Analyze transparently as exploratory |
|
Peripheral
|
Interesting but not needed for this study |
Exclude or reserve for future research |
This forces the study to establish priorities before every interesting construct finds its way into the questionnaire.
The literature should help reduce variables, not only generate them
A good literature review does more than produce a long list of candidate predictors.
It should also help researchers reject weak possibilities.
Some relationships may have inconsistent evidence. Some constructs may overlap strongly. Some variables may belong to neighboring questions rather than the present one.
Previous research can therefore help narrow the model as much as expand it.
This is another reason not to include a variable simply because previous studies did.