03 · What You Need to Know
“Controlled For” Describes an Action, Not a Causal Justification
What is a control variable?
The phrase control variable is used broadly across research traditions. In regression-based studies, it often refers to a variable included in a model so that the focal X–Y relationship is estimated conditional on that variable.
Suppose researchers model academic performance as a function of educational technology use while also including age and prior achievement in the regression equation. They may describe age and prior achievement as control variables because the coefficient for technology use is estimated while holding those variables constant in the statistical model.
This terminology tells us what happened analytically. It does not tell us whether age or prior achievement needed to be adjusted for to estimate a causal effect.
Depending on the research question, a control variable might be:
- a genuine confounder;
- a precision variable;
- a mediator;
- a collider or descendant of a collider;
- a proxy for another relevant variable;
- a baseline predictor included for another design or modeling reason.
These variables should not all be treated as causally equivalent merely because they appear in the same regression equation.
What is a confounder?
Confounding arises when the exposed and unexposed, treated and untreated, or otherwise compared groups differ in ways that also influence the outcome, creating a noncausal component in the observed exposure–outcome association.
In a simple causal diagram, a common cause C of exposure X and outcome Y creates a backdoor path:
X ← C → Y
If the goal is to estimate the causal effect of X on Y, that open backdoor path can bias a simple comparison between different values of X.
For example, suppose researchers examine whether voluntary attendance at supplemental tutorials improves examination performance. Prior academic achievement may affect students' likelihood of attending tutorials and also affect later examination scores:
Tutorial attendance ← Prior achievement → Examination performance
If so, students who attend and do not attend tutorials may differ systematically before tutorial participation even begins. Part of the observed association may therefore reflect prior achievement rather than the causal effect of tutorial attendance.
Confounder
A causal role relevant to bias in estimating a particular exposure–outcome effect.
Control variable
A variable conditioned on or adjusted for in the research design or statistical analysis.
A confounder is defined relative to a particular causal question
Age, sex, socioeconomic status, prior achievement, and similar variables are frequently described as “known confounders.” This shorthand can be useful in a specific substantive context, but no variable is universally a confounder for every analysis.
The role depends on:
- the exposure;
- the outcome;
- the population;
- the timing of the variables;
- the causal effect or estimand of interest;
- the assumed causal structure.
Prior achievement might confound an observational comparison between voluntary tutoring and later examination performance. It would not automatically confound every relationship in an educational dataset.
This is why researchers should be cautious about importing the adjustment set from another article without asking whether the earlier study addressed the same causal question.
Association with X and Y is not enough to define a confounder
Older and purely data-driven approaches sometimes identify potential confounders by checking whether a variable is statistically associated with both exposure and outcome or whether adding it changes the focal coefficient by a chosen percentage.
Such criteria can be misleading because mediators, colliders, and other variables can also display statistical associations with X and Y.
Consider:
X → M → Y
The mediator M may be strongly associated with both X and Y, yet if the target is the total causal effect of X on Y, adjusting for M can remove part of the causal pathway that contributes to that total effect.
The distinction among a confounder, mediator, and moderator therefore cannot be recovered from a correlation matrix alone.
Why adjusting for a genuine confounder can help
When an appropriate set of measured covariates is sufficient to block relevant noncausal backdoor paths between X and Y, conditioning on that set can help make compared exposure groups exchangeable with respect to those measured causes.
Regression adjustment is one possible strategy. Others include stratification, standardization, matching, propensity-score methods, inverse-probability weighting, restriction, and design-based approaches.
The objective is not statistical control for its own sake. It is to obtain a comparison that more closely corresponds to the causal effect of interest under the required assumptions.
Watch Out
Adjustment for measured confounders does not establish that all confounding has been removed. A causal interpretation still depends on assumptions about unmeasured confounding, measurement, selection, positivity, model specification, and other features of the study.
Controlling for a mediator can change the question you are answering
Suppose researchers investigate whether an educational intervention improves achievement partly because it increases student engagement:
Intervention → Engagement → Achievement
If the researchers want the total effect of the intervention, routinely adjusting for engagement blocks part of the pathway through which the intervention is proposed to work.
The resulting coefficient no longer represents the same total-effect question.
If the researchers instead want to distinguish direct and indirect pathways, engagement may need to be modeled explicitly as a mediator under an appropriate mediation framework.
Calling engagement “a control” obscures this distinction. Whether adjustment is appropriate depends on what effect the study intends to estimate and what it means for a mediating variable to lie on the X–Y pathway.
Controlling for a collider can create bias
One of the most counterintuitive lessons of causal inference is that conditioning on a variable can create an association rather than remove one.
Consider a structure in which X and Y both cause C:
X → C ← Y
C is called a collider because two arrows collide at it. Without conditioning on C, this pathway is naturally closed. Conditioning on C can open a noncausal pathway between X and Y.
A familiar intuitive example involves selection. Suppose both academic ability and extraordinary extracurricular achievement increase the chance of admission to a highly selective program. Among admitted students only, lower values of one characteristic may become statistically associated with higher values of the other because admission depends on both.
The important point is broader: adjustment is not automatically protective against bias. Depending on the causal structure, it can introduce bias.
This is why “adjust for everything available” is not a safe rule
Imagine a dataset containing 30 variables and a researcher who enters all 30 into a regression model “to control for possible confounding.” The approach may feel rigorous because many alternative explanations have apparently been addressed.
But causal adjustment is not a contest to maximize the number of covariates.
Among those 30 variables may be:
- appropriate confounders;
- mediators on the causal pathway;
- colliders;
- variables caused by the exposure;
- irrelevant variables;
- highly correlated proxies;
- variables measured after the outcome.
Adjusting indiscriminately can therefore change the estimand, amplify measurement problems, reduce precision, or introduce bias.
More variables can consequently make an analysis weaker rather than stronger, which is why the question of whether adding more variables can weaken a study has no simple “more is better” answer.
A sufficient adjustment set is more important than a long adjustment set
Modern causal-inference approaches often focus on identifying a sufficient adjustment set: a set of variables that, under the assumed causal structure, blocks the relevant noncausal backdoor paths without conditioning on inappropriate variables.
There can sometimes be more than one sufficient adjustment set.
This means researchers need not necessarily adjust for every ancestor or correlate of X and Y. The objective is to choose variables that are sufficient for the causal identification strategy while avoiding variables that alter or bias the causal comparison.
The phrase “we adjusted for all available covariates” is therefore not inherently reassuring. A clear rationale for the adjustment set is usually more informative.
Directed acyclic graphs can help clarify adjustment decisions
A directed acyclic graph, or DAG, represents assumed causal relationships using nodes and directed arrows.
DAGs can help researchers distinguish:
- common causes that generate confounding;
- mediators lying on causal pathways;
- colliders on paths that should generally remain closed;
- alternative sufficient adjustment sets.
A DAG does not discover the causal structure from the data. Researchers build it using substantive knowledge and assumptions.
That limitation is also its strength: the diagram makes assumptions visible. Instead of hiding variable-selection choices inside a statistical model, researchers must state why they believe one variable precedes another and why a particular path should or should not be blocked.
You may control for a variable even when it is not a confounder
Not every nonconfounder adjustment is necessarily wrong.
In randomized experiments, for example, baseline variables may be included to improve precision even though randomization means those variables are not confounders in the same sense as in a poorly controlled observational comparison.
Researchers may also adjust for design factors such as stratification variables, blocking variables, site indicators, or baseline measurements for reasons tied to the analysis plan.
The key is to describe the reason accurately.
| Reason for including a variable |
Does that automatically make it a confounder? |
Possible purpose |
| Blocks a relevant backdoor path |
Potentially, as part of a sufficient confounding adjustment set |
Reduce confounding bias |
| Strongly predicts the outcome in a randomized study |
No |
Improve precision |
| Lies on the X → Y pathway |
No |
Mediation or direct-effect analysis |
| Included because previous papers used it |
No |
Requires independent justification |
| Interacts with X |
No |
Investigate effect heterogeneity |
Control variable and covariate are also not synonymous with confounder
The word covariate is often used broadly for variables included alongside an exposure or predictor in a statistical model. This may include confounders, precision variables, baseline measurements, moderators, or other predictors.
Thus:
confounder describes a causal role;
covariate commonly describes a variable's presence or function in an analytical model;
control variable commonly emphasizes that the variable is being conditioned on while estimating another relationship.
The distinctions are not perfectly standardized across every discipline, which makes it especially important to state what variables were adjusted for and why. The relationship between a covariate and a control variable is therefore partly terminological and partly methodological.
Baseline status does not automatically make something a confounder
Researchers sometimes assume that every pre-exposure variable is safe or necessary to control.
Being measured before X prevents a variable from being a consequence of X, which can be important. But temporal precedence alone does not make it a confounder.
A baseline variable may be unrelated to the causal pathways creating confounding. Another baseline variable may be an instrument-like cause of X that does not itself cause Y. Depending on the structure and measurement, adjusting for such variables may provide little benefit and can sometimes have undesirable statistical consequences.
Pre-exposure timing is therefore useful information, not a complete variable-selection rule.
Statistical significance is a poor rule for selecting confounders
A variable does not stop being part of a relevant confounding structure because its coefficient has p >.05 in one sample.
Conversely, statistical significance does not transform an ordinary predictor into a confounder.
Confounding concerns bias in a causal comparison. Variable selection for confounding control should therefore be based primarily on substantive causal knowledge and the identification strategy, not on stepwise p-value filtering.
Data-driven procedures may remove an important confounder because its sample association happens to be imprecise or retain an inappropriate variable because it strongly predicts the outcome.
Change-in-estimate rules do not determine causal roles by themselves
Another common strategy is to add a candidate variable and classify it as a confounder if the focal coefficient changes by a specified percentage, such as 10%.
A change in estimate can be diagnostically interesting, but it does not reveal the variable's causal role.
A mediator can change the coefficient substantially. So can a collider, measurement error, noncollapsibility in some models, or changes in the analyzed sample due to missing data.
Causal reasoning should therefore precede mechanical change-in-estimate rules.
Whether to adjust can depend on the estimand
The correct adjustment strategy depends on what effect the researcher wants.
Suppose X affects M and both X and M affect Y:
X → M → Y
X → Y
If the target is the total effect of X on Y, blocking M is generally inconsistent with that target because the mediated pathway is part of the total effect.
If the target is a particular direct effect, the role of M becomes different, and appropriate causal mediation methods and additional assumptions may be required.
Thus, asking “Should I control for M?” before specifying the estimand puts the analysis in the wrong order. The causal question comes first.
The same variable can be a confounder in one study and not another
Suppose digital literacy influences whether students voluntarily adopt an AI study tool and also affects their learning outcomes. Digital literacy may confound the tool-use–achievement relationship.
Now consider a randomized trial in which the tool is assigned independently of digital literacy. Digital literacy can no longer cause assignment because assignment was randomized, although it may still predict the outcome or modify the intervention effect.
The construct remains the same; its role changes because the exposure-generating process changes.
Similarly, a variable may occupy a different causal position when the research question changes. This is why the same variable can be a confounder in one study and a mediator in another.
Previous studies are evidence, not an automatic adjustment recipe
A literature review can help identify plausible common causes and support the construction of a causal model. But copying every covariate from an influential previous paper is not a substitute for causal reasoning.
The previous study may have:
- studied a different exposure or outcome;
- estimated a different causal effect;
- used the variable for precision rather than confounding;
- followed an older adjustment convention;
- made assumptions that do not apply to your setting.
The fact that previous studies included a variable is therefore a reason to investigate its role, not an instruction to reproduce the same model automatically.
“Controlled for” should not be used as shorthand for “all alternative explanations removed”
Even a carefully chosen adjustment set cannot guarantee that an observational estimate is free from bias.
Potential problems may include unmeasured confounding, poorly measured confounders, selection bias, missing data, model misspecification, positivity violations, interference, and inaccurate assumptions about the causal structure.
Researchers should therefore avoid statements implying that adjustment has eliminated every alternative explanation.
A more defensible description is usually specific: identify the variables or adjustment set used, explain the rationale, and acknowledge the assumptions under which the adjusted estimate receives a causal interpretation.