Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Control Variable vs. Confounder: Are They the Same Thing?

A control variable is a variable a researcher holds constant or adjusts for analytically, whereas a confounder has a specific causal role that can bias an exposure–outcome comparison. Not every control variable is a confounder, and not every available variable should be controlled.

105
Control Variable vs. Confounder Guide 105 of 223
01 · The Question

If You “Control for” a Variable, Does That Mean It Was a Confounder?

Research articles routinely state that analyses “controlled for age, sex, socioeconomic status, prior achievement, and other covariates.” These variables are sometimes called control variables and sometimes confounders, often as though the terms were interchangeable.

They are not.

A control variable is defined largely by what the researcher does with it: the variable is held constant, conditioned on, stratified by, matched on, or otherwise adjusted for in the design or analysis. A confounder, by contrast, has a particular role in a causal structure. It contributes to a noncausal association between an exposure and outcome and therefore threatens estimation of the causal effect of interest.

The distinction matters because “controlling for more variables” does not automatically produce a more trustworthy estimate. Adjusting for an appropriate set of confounders may reduce bias. Adjusting for a mediator can remove part of the causal effect you intended to estimate. Adjusting for a collider can create bias that was absent before adjustment.

The critical question is therefore not simply “What did you control for?” but “Why should this variable be controlled for in relation to the causal effect you are trying to estimate?”

02 · The Short Answer

A Confounder Is a Causal Role; a Control Variable Is an Analytical Role

In Brief

A confounder is a variable or causal structure that can bias the estimated causal relationship between an exposure and outcome, whereas a control variable is any variable that a researcher conditions on or adjusts for in the design or analysis.

A confounder may appropriately be used as a control variable, but the categories are not equivalent. Researchers can control for variables that are not confounders, and some such adjustments can reduce precision, change the causal effect being estimated, or even introduce bias.

03 · What You Need to Know

“Controlled For” Describes an Action, Not a Causal Justification

What is a control variable?

The phrase control variable is used broadly across research traditions. In regression-based studies, it often refers to a variable included in a model so that the focal X–Y relationship is estimated conditional on that variable.

Suppose researchers model academic performance as a function of educational technology use while also including age and prior achievement in the regression equation. They may describe age and prior achievement as control variables because the coefficient for technology use is estimated while holding those variables constant in the statistical model.

This terminology tells us what happened analytically. It does not tell us whether age or prior achievement needed to be adjusted for to estimate a causal effect.

Depending on the research question, a control variable might be:

  • a genuine confounder;
  • a precision variable;
  • a mediator;
  • a collider or descendant of a collider;
  • a proxy for another relevant variable;
  • a baseline predictor included for another design or modeling reason.

These variables should not all be treated as causally equivalent merely because they appear in the same regression equation.

What is a confounder?

Confounding arises when the exposed and unexposed, treated and untreated, or otherwise compared groups differ in ways that also influence the outcome, creating a noncausal component in the observed exposure–outcome association.

In a simple causal diagram, a common cause C of exposure X and outcome Y creates a backdoor path:

X ← C → Y

If the goal is to estimate the causal effect of X on Y, that open backdoor path can bias a simple comparison between different values of X.

For example, suppose researchers examine whether voluntary attendance at supplemental tutorials improves examination performance. Prior academic achievement may affect students' likelihood of attending tutorials and also affect later examination scores:

Tutorial attendance ← Prior achievement → Examination performance

If so, students who attend and do not attend tutorials may differ systematically before tutorial participation even begins. Part of the observed association may therefore reflect prior achievement rather than the causal effect of tutorial attendance.

Confounder A causal role relevant to bias in estimating a particular exposure–outcome effect.
Control variable A variable conditioned on or adjusted for in the research design or statistical analysis.

A confounder is defined relative to a particular causal question

Age, sex, socioeconomic status, prior achievement, and similar variables are frequently described as “known confounders.” This shorthand can be useful in a specific substantive context, but no variable is universally a confounder for every analysis.

The role depends on:

  • the exposure;
  • the outcome;
  • the population;
  • the timing of the variables;
  • the causal effect or estimand of interest;
  • the assumed causal structure.

Prior achievement might confound an observational comparison between voluntary tutoring and later examination performance. It would not automatically confound every relationship in an educational dataset.

This is why researchers should be cautious about importing the adjustment set from another article without asking whether the earlier study addressed the same causal question.

Association with X and Y is not enough to define a confounder

Older and purely data-driven approaches sometimes identify potential confounders by checking whether a variable is statistically associated with both exposure and outcome or whether adding it changes the focal coefficient by a chosen percentage.

Such criteria can be misleading because mediators, colliders, and other variables can also display statistical associations with X and Y.

Consider:

X → M → Y

The mediator M may be strongly associated with both X and Y, yet if the target is the total causal effect of X on Y, adjusting for M can remove part of the causal pathway that contributes to that total effect.

The distinction among a confounder, mediator, and moderator therefore cannot be recovered from a correlation matrix alone.

Why adjusting for a genuine confounder can help

When an appropriate set of measured covariates is sufficient to block relevant noncausal backdoor paths between X and Y, conditioning on that set can help make compared exposure groups exchangeable with respect to those measured causes.

Regression adjustment is one possible strategy. Others include stratification, standardization, matching, propensity-score methods, inverse-probability weighting, restriction, and design-based approaches.

The objective is not statistical control for its own sake. It is to obtain a comparison that more closely corresponds to the causal effect of interest under the required assumptions.

Watch Out

Adjustment for measured confounders does not establish that all confounding has been removed. A causal interpretation still depends on assumptions about unmeasured confounding, measurement, selection, positivity, model specification, and other features of the study.

Controlling for a mediator can change the question you are answering

Suppose researchers investigate whether an educational intervention improves achievement partly because it increases student engagement:

Intervention → Engagement → Achievement

If the researchers want the total effect of the intervention, routinely adjusting for engagement blocks part of the pathway through which the intervention is proposed to work.

The resulting coefficient no longer represents the same total-effect question.

If the researchers instead want to distinguish direct and indirect pathways, engagement may need to be modeled explicitly as a mediator under an appropriate mediation framework.

Calling engagement “a control” obscures this distinction. Whether adjustment is appropriate depends on what effect the study intends to estimate and what it means for a mediating variable to lie on the X–Y pathway.

Controlling for a collider can create bias

One of the most counterintuitive lessons of causal inference is that conditioning on a variable can create an association rather than remove one.

Consider a structure in which X and Y both cause C:

X → C ← Y

C is called a collider because two arrows collide at it. Without conditioning on C, this pathway is naturally closed. Conditioning on C can open a noncausal pathway between X and Y.

A familiar intuitive example involves selection. Suppose both academic ability and extraordinary extracurricular achievement increase the chance of admission to a highly selective program. Among admitted students only, lower values of one characteristic may become statistically associated with higher values of the other because admission depends on both.

The important point is broader: adjustment is not automatically protective against bias. Depending on the causal structure, it can introduce bias.

This is why “adjust for everything available” is not a safe rule

Imagine a dataset containing 30 variables and a researcher who enters all 30 into a regression model “to control for possible confounding.” The approach may feel rigorous because many alternative explanations have apparently been addressed.

But causal adjustment is not a contest to maximize the number of covariates.

Among those 30 variables may be:

  • appropriate confounders;
  • mediators on the causal pathway;
  • colliders;
  • variables caused by the exposure;
  • irrelevant variables;
  • highly correlated proxies;
  • variables measured after the outcome.

Adjusting indiscriminately can therefore change the estimand, amplify measurement problems, reduce precision, or introduce bias.

More variables can consequently make an analysis weaker rather than stronger, which is why the question of whether adding more variables can weaken a study has no simple “more is better” answer.

A sufficient adjustment set is more important than a long adjustment set

Modern causal-inference approaches often focus on identifying a sufficient adjustment set: a set of variables that, under the assumed causal structure, blocks the relevant noncausal backdoor paths without conditioning on inappropriate variables.

There can sometimes be more than one sufficient adjustment set.

This means researchers need not necessarily adjust for every ancestor or correlate of X and Y. The objective is to choose variables that are sufficient for the causal identification strategy while avoiding variables that alter or bias the causal comparison.

The phrase “we adjusted for all available covariates” is therefore not inherently reassuring. A clear rationale for the adjustment set is usually more informative.

Directed acyclic graphs can help clarify adjustment decisions

A directed acyclic graph, or DAG, represents assumed causal relationships using nodes and directed arrows.

DAGs can help researchers distinguish:

  • common causes that generate confounding;
  • mediators lying on causal pathways;
  • colliders on paths that should generally remain closed;
  • alternative sufficient adjustment sets.

A DAG does not discover the causal structure from the data. Researchers build it using substantive knowledge and assumptions.

That limitation is also its strength: the diagram makes assumptions visible. Instead of hiding variable-selection choices inside a statistical model, researchers must state why they believe one variable precedes another and why a particular path should or should not be blocked.

You may control for a variable even when it is not a confounder

Not every nonconfounder adjustment is necessarily wrong.

In randomized experiments, for example, baseline variables may be included to improve precision even though randomization means those variables are not confounders in the same sense as in a poorly controlled observational comparison.

Researchers may also adjust for design factors such as stratification variables, blocking variables, site indicators, or baseline measurements for reasons tied to the analysis plan.

The key is to describe the reason accurately.

Reason for including a variable Does that automatically make it a confounder? Possible purpose
Blocks a relevant backdoor path Potentially, as part of a sufficient confounding adjustment set Reduce confounding bias
Strongly predicts the outcome in a randomized study No Improve precision
Lies on the X → Y pathway No Mediation or direct-effect analysis
Included because previous papers used it No Requires independent justification
Interacts with X No Investigate effect heterogeneity

Control variable and covariate are also not synonymous with confounder

The word covariate is often used broadly for variables included alongside an exposure or predictor in a statistical model. This may include confounders, precision variables, baseline measurements, moderators, or other predictors.

Thus:

confounder describes a causal role;

covariate commonly describes a variable's presence or function in an analytical model;

control variable commonly emphasizes that the variable is being conditioned on while estimating another relationship.

The distinctions are not perfectly standardized across every discipline, which makes it especially important to state what variables were adjusted for and why. The relationship between a covariate and a control variable is therefore partly terminological and partly methodological.

Baseline status does not automatically make something a confounder

Researchers sometimes assume that every pre-exposure variable is safe or necessary to control.

Being measured before X prevents a variable from being a consequence of X, which can be important. But temporal precedence alone does not make it a confounder.

A baseline variable may be unrelated to the causal pathways creating confounding. Another baseline variable may be an instrument-like cause of X that does not itself cause Y. Depending on the structure and measurement, adjusting for such variables may provide little benefit and can sometimes have undesirable statistical consequences.

Pre-exposure timing is therefore useful information, not a complete variable-selection rule.

Statistical significance is a poor rule for selecting confounders

A variable does not stop being part of a relevant confounding structure because its coefficient has p >.05 in one sample.

Conversely, statistical significance does not transform an ordinary predictor into a confounder.

Confounding concerns bias in a causal comparison. Variable selection for confounding control should therefore be based primarily on substantive causal knowledge and the identification strategy, not on stepwise p-value filtering.

Data-driven procedures may remove an important confounder because its sample association happens to be imprecise or retain an inappropriate variable because it strongly predicts the outcome.

Change-in-estimate rules do not determine causal roles by themselves

Another common strategy is to add a candidate variable and classify it as a confounder if the focal coefficient changes by a specified percentage, such as 10%.

A change in estimate can be diagnostically interesting, but it does not reveal the variable's causal role.

A mediator can change the coefficient substantially. So can a collider, measurement error, noncollapsibility in some models, or changes in the analyzed sample due to missing data.

Causal reasoning should therefore precede mechanical change-in-estimate rules.

Whether to adjust can depend on the estimand

The correct adjustment strategy depends on what effect the researcher wants.

Suppose X affects M and both X and M affect Y:

X → M → Y
X → Y

If the target is the total effect of X on Y, blocking M is generally inconsistent with that target because the mediated pathway is part of the total effect.

If the target is a particular direct effect, the role of M becomes different, and appropriate causal mediation methods and additional assumptions may be required.

Thus, asking “Should I control for M?” before specifying the estimand puts the analysis in the wrong order. The causal question comes first.

The same variable can be a confounder in one study and not another

Suppose digital literacy influences whether students voluntarily adopt an AI study tool and also affects their learning outcomes. Digital literacy may confound the tool-use–achievement relationship.

Now consider a randomized trial in which the tool is assigned independently of digital literacy. Digital literacy can no longer cause assignment because assignment was randomized, although it may still predict the outcome or modify the intervention effect.

The construct remains the same; its role changes because the exposure-generating process changes.

Similarly, a variable may occupy a different causal position when the research question changes. This is why the same variable can be a confounder in one study and a mediator in another.

Previous studies are evidence, not an automatic adjustment recipe

A literature review can help identify plausible common causes and support the construction of a causal model. But copying every covariate from an influential previous paper is not a substitute for causal reasoning.

The previous study may have:

  • studied a different exposure or outcome;
  • estimated a different causal effect;
  • used the variable for precision rather than confounding;
  • followed an older adjustment convention;
  • made assumptions that do not apply to your setting.

The fact that previous studies included a variable is therefore a reason to investigate its role, not an instruction to reproduce the same model automatically.

“Controlled for” should not be used as shorthand for “all alternative explanations removed”

Even a carefully chosen adjustment set cannot guarantee that an observational estimate is free from bias.

Potential problems may include unmeasured confounding, poorly measured confounders, selection bias, missing data, model misspecification, positivity violations, interference, and inaccurate assumptions about the causal structure.

Researchers should therefore avoid statements implying that adjustment has eliminated every alternative explanation.

A more defensible description is usually specific: identify the variables or adjustment set used, explain the rationale, and acknowledge the assumptions under which the adjusted estimate receives a causal interpretation.

04 · A Practical Example

Why Three Plausible “Controls” Should Not Automatically Be Treated the Same Way

Hypothetical Example

Voluntary AI tutor use and examination performance

Researchers use observational data to investigate whether voluntary use of an AI tutoring platform improves students' final examination scores. They have measurements of prior achievement, time spent practicing after beginning to use the platform, and participation in an advanced academic support program.

Prior achievement Suppose stronger students are more likely to adopt the AI tutor and also tend to obtain higher final scores. Prior achievement may lie on a backdoor path and could form part of the confounding adjustment set.
Practice time Suppose using the AI tutor encourages students to practice more, which then improves examination performance. Practice time is downstream of the exposure and may be a mediator. Adjusting for it would change a total-effect question.
Academic support participation Suppose both struggling performance during the semester and intensive AI-tool use increase the probability that students are referred to the support program. Conditioning on program participation could create a selection or collider problem depending on the causal structure.
Decision The researchers should not place all three variables into a model merely because each is available. They should first specify the causal effect of interest and determine which variables belong in an appropriate adjustment set.

The example illustrates why the phrase “we controlled for several relevant factors” conveys too little information. What matters is the causal role of each variable and the effect being estimated.

Prior achievement may appropriately be controlled to address confounding. Practice time may instead be part of the mechanism. Conditioning on a selection variable could even introduce bias. The same statistical operation, adding a variable to a regression model, can therefore have very different consequences.

05 · What Researchers Often Get Wrong

Common Misconceptions About Statistical Control and Confounding

Misconception

Every control variable is a confounder

No. Researchers control for variables for many reasons, including precision, design adjustment, mediation questions, subgroup modeling, or convention. Confounder is a specific causal role, not a synonym for every variable included alongside X in a model.

Misconception

If a variable predicts the outcome, you should control for it

Outcome prediction alone does not determine whether adjustment is appropriate for causal inference. A strong outcome predictor may be useful for precision, but another outcome predictor could be a mediator or collider whose adjustment changes or biases the causal estimate.

Misconception

Controlling for more variables always reduces bias

No. Appropriate confounding adjustment can reduce bias, but adjusting for variables on causal pathways or for colliders can change the estimand or introduce bias. Covariate quantity is not a measure of causal rigor.

Misconception

If Z changes the coefficient for X, Z must be a confounder

A coefficient can change for many reasons. Mediators, colliders, model nonlinearity, noncollapsibility, missing-data changes, and other factors can alter estimates. The causal role of Z requires substantive justification rather than a mechanical change-in-estimate rule.

Misconception

A nonsignificant covariate can be safely removed from a causal model

Statistical significance does not determine whether a variable is needed to address confounding. A causally important variable can have an imprecisely estimated coefficient in a particular sample while still being necessary as part of the adjustment strategy.

Misconception

Any variable measured before the exposure is safe to control

Pre-exposure timing rules out some problematic roles, such as being caused by the exposure, but does not automatically make the variable necessary or beneficial for adjustment. Its causal connections to the exposure and outcome still matter.

Misconception

After controlling for confounders, an observational estimate is proven causal

No. Causal interpretation still rests on assumptions, including adequate control of relevant confounding and the absence of other consequential sources of bias. Statistical adjustment cannot verify all of those assumptions from the observed data alone.

06 · What This Means for You

Choose Adjustment Variables From the Causal Question, Not From the Regression Output

Before deciding what to control for, define the relationship and effect you actually want to estimate. Then determine what causal structure would make that effect identifiable from your data.

A simple decision framework

If the variable is part of a sufficient set needed to block noncausal backdoor paths
Adjusting for it may be appropriate for confounding control, subject to the assumptions of the causal model.
If the variable is caused by X and lies on the pathway to Y
Do not automatically control for it when estimating the total effect; determine whether the research question instead concerns mediation or a direct effect.
If the variable is a collider on an otherwise closed path
Conditioning on it can introduce bias, so adjustment may be harmful.
If you are including the variable only because previous studies did
Reassess its role in your own exposure, outcome, timing, population, and causal model before reproducing the adjustment.

This reasoning should occur as part of deciding which variables actually belong in the study, ideally before the final analytical model is fitted.

A useful methods section should also explain the basis for adjustment. Instead of simply stating that “demographic variables were controlled,” identify why particular variables were selected and, for a causal analysis, how they relate to the assumed causal structure or adjustment strategy.

The goal is not to demonstrate that you controlled for many things. It is to show that the model corresponds to a clearly defined research question.

07 · A Quick Checklist

Before Adding a “Control Variable,” Check Why It Belongs There

Before adjusting for a variable, check:
Define the exposure, outcome, population, and causal effect or estimand of interest.
Identify the causal reason the candidate variable might need to be adjusted for.
Distinguish genuine confounding control from adjustment for precision or other analytical purposes.
Check whether the candidate variable is downstream of the exposure and could be a mediator.
Consider whether conditioning on the variable could open a collider or selection pathway.
Do not select confounders solely from p-values, correlations, or change-in-estimate thresholds.
Use substantive knowledge and, when helpful, a causal diagram to justify the adjustment set.
Do not copy adjustment variables from previous studies without checking whether their causal roles apply to your research question.
Report what was adjusted for and why rather than implying that statistical control removed every alternative explanation.
08 · Frequently Asked Questions

Frequently Asked Questions About Control Variables and Confounders

Is every confounder a control variable?

A measured confounder may be addressed through statistical adjustment and therefore function as a control variable in a particular analysis, but confounding can also be handled through design or other analytical strategies. More importantly, “confounder” describes the causal problem, whereas “control variable” describes a way a variable is treated analytically.

Is every control variable a confounder?

No. Researchers may adjust for baseline predictors, design variables, precision variables, mediators, or other covariates. Whether those adjustments are appropriate depends on the research objective and causal structure.

How do I know whether a variable is a confounder?

Start with the causal question and consider whether the variable participates in a noncausal pathway between the exposure and outcome that needs to be blocked. Subject-matter knowledge, temporal ordering, prior evidence, and causal diagrams can help identify an appropriate adjustment set.

Should I control for anything associated with the outcome?

No. Outcome association alone does not establish that adjustment is necessary. Some strongly outcome-related variables may improve precision, while others may lie on a causal pathway or introduce bias if conditioned on.

Should I remove a control variable if it is not statistically significant?

Not automatically. A variable required for a causal adjustment strategy does not become unnecessary because its individual coefficient is nonsignificant in one sample. Variable selection should follow the purpose of the model rather than a mechanical significance threshold.

Can controlling for a variable make bias worse?

Yes. Conditioning on a collider can open a noncausal pathway, and controlling for a mediator can remove part of a total causal effect. Adjustment should therefore be based on the assumed causal structure rather than the principle that more controls are always safer.

Is a covariate the same as a control variable?

The terms overlap substantially in ordinary statistical writing, but usage varies. Covariate is often a broad term for a variable included in a model, whereas control variable usually emphasizes that its value is being conditioned on while another relationship is estimated. Neither term automatically means confounder.

Does controlling for confounders prove causality?

No. Confounding adjustment can support causal inference when an appropriate set of variables is measured and correctly handled, but causal interpretation still depends on assumptions about unmeasured confounding, measurement, selection, model specification, and other features of the study.

09 · The Bottom Line

Do Not Call Every Variable You Adjust for a Confounder

The Bottom Line

A confounder has a specific role in creating bias in a causal exposure–outcome comparison, while a control variable is simply a variable that the researcher conditions on or adjusts for in the design or analysis.

Appropriate adjustment depends on the causal question and the role of each variable. Controlling for genuine confounding may reduce bias, but indiscriminate adjustment can change the effect being estimated or even introduce new bias. The strongest justification is therefore not “we controlled for many variables,” but “these variables were selected because they are appropriate for this estimand and causal structure.”

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes