03 · What You Need to Know
“Effect” Is Primarily an Inferential Claim, Not a Design Label
In ordinary conversation, “effect” can simply mean a result or consequence. In causal research, however, the term has a more specific implication: an effect concerns how an outcome would differ under alternative exposure or intervention conditions.
That is why methodological guidance on causal inference treats “effect” as causal language. If the statistical quantity estimated by a study cannot be interpreted as a causal effect, researchers can instead describe the association between the exposure and outcome.
This distinction is more useful than asking whether the study is labeled “experimental” or “observational.” The relevant questions are: What exactly are you trying to estimate? What comparison defines the effect? What design generates the evidence? Which assumptions allow the observed data to identify the causal quantity of interest? And are those assumptions credible?
Association and effect answer different questions
Consider two questions:
“Is generative AI use associated with academic writing performance among undergraduate students?”
“What is the effect of access to generative AI on academic writing performance among undergraduate students?”
The first asks whether two observed characteristics are related. The second asks a counterfactual question: how would writing performance differ under alternative conditions of AI access?
Association
How does the observed outcome differ across observed values or categories of an exposure or predictor?
Causal effect
How would the outcome differ under alternative exposure or intervention conditions for a defined target population?
An association can exist without representing a causal effect. Students who choose to use AI may differ from students who do not in prior writing ability, digital literacy, motivation, workload, socioeconomic circumstances, course requirements, or many other characteristics. Those differences can complicate causal interpretation.
Randomized experiments have a major advantage for causal inference
Randomization is powerful because, when properly implemented with sufficient sample size and appropriate analysis, treatment assignment helps create groups that are comparable with respect to both measured and unmeasured baseline characteristics in expectation.
This can make the causal comparison much more credible than simply comparing people who selected different exposures themselves.
For that reason, randomized controlled trials occupy an important position in causal research. But randomization is a method for supporting causal inference, not the definition of causality itself.
Some scientifically important exposures cannot ethically or practically be randomized. Researchers do not randomly assign people to smoke cigarettes for decades, experience air pollution, live in poverty, or suffer traumatic events merely to make causal inference convenient. In such settings, causal questions may still matter enormously.
Observational studies can be designed around causal questions
Contemporary causal-inference methods explicitly address estimation of causal effects using observational data. Methodological frameworks based on potential outcomes, structural causal models, directed acyclic graphs, target trials, natural experiments, instrumental variables, inverse probability weighting, and related approaches have been developed to make the causal question and its assumptions explicit.
Recent guidance for medical journals likewise argues that observational studies may aim to provide evidence about causal effects. It recommends evaluating such studies by asking what the causal question is, what quantity would answer it, what design is used, what causal assumptions are required, how the observed data identify that quantity, and whether a causal interpretation is tenable.
So “observational” and “causal” are not mutually exclusive categories.
But ordinary observational analysis does not automatically estimate an effect
This qualification is essential.
Suppose you survey students once and measure self-reported generative AI use and writing scores. You then run a regression in which writing performance is the outcome, AI use is the predictor, and several demographic characteristics are entered as covariates.
Calling the regression coefficient an “effect” does not make it causal.
You would need to consider whether the exposure is well defined, whether relevant confounding has been addressed, whether temporal ordering is appropriate, whether selection and measurement processes introduce bias, whether the statistical model corresponds to the intended causal estimand, and whether the assumptions required for identification are defensible.
Causal-effect estimation from observational data commonly relies on assumptions concerning exchangeability, positivity, consistency, and related conditions. Confounding, selection bias, measurement bias, collider bias, and model misspecification can undermine causal interpretation.
Watch Out
“Adjusted for age, sex, and several other variables” is not synonymous with “causal.” Statistical adjustment supports causal inference only when the adjustment strategy follows an appropriate causal model and the assumptions needed for the intended effect estimate are sufficiently credible.
You need to define what intervention or exposure contrast you mean
“What is the effect of social media?” is not yet a well-defined causal question.
Effect compared with what? No social media use? Thirty minutes less use per day? Removing one platform? Preventing use during study periods? Changing from active posting to passive browsing?
Causal effects are contrasts between specified conditions. The more ambiguous the exposure, the harder it becomes to say what intervention or alternative state the causal estimate represents.
This problem appears in educational technology research as well. “What is the effect of AI use on learning?” could refer to access to an AI tool, frequency of use, particular instructional uses, particular prompts, AI-generated feedback, or an institutional intervention encouraging AI-assisted learning.
A causal question becomes more informative when the relevant contrast can be stated clearly enough that the hypothetical alternatives are understandable.
The outcome and time horizon also matter
An effect is not simply “the effect of X.” It is the effect of X on a particular outcome over a relevant period.
Generative AI might improve immediate task performance while having a different relationship with later independent performance. An educational intervention might improve examination scores at the end of a course but show little difference one year later.
Therefore, an effect question should usually be specific enough about the outcome and, when substantively important, the time horizon. This is part of determining how specific the research question should be before the study begins.
Confounding is central to observational causal inference
A confounder is, broadly, a factor that creates difficulty in interpreting an observed exposure-outcome relationship causally because it is related to both the exposure and outcome in the relevant causal structure.
Suppose students with stronger prior writing skills are more comfortable experimenting with AI and also tend to achieve higher subsequent writing scores. A simple comparison between AI users and nonusers could partly reflect prior writing ability rather than an effect of AI use.
Alternatively, students struggling with writing might use AI more often. In that case, the observed relationship could even run in the opposite direction from the causal effect of interest.
Causal inference requires researchers to reason about these structures before interpreting adjusted estimates. Directed acyclic graphs and other causal models can help make assumptions about relationships among variables explicit.
More covariates are not necessarily better
A common response to confounding is to adjust for every available variable. That can create new problems.
Some variables may lie on the causal pathway between exposure and outcome. Others may be colliders whose conditioning can introduce bias. Still others may be irrelevant to the causal identification problem.
The decision about what to adjust for should therefore follow causal reasoning rather than a software-generated list of statistically significant predictors. Causal-effect estimation is not simply multiple regression with a larger covariate table.
Temporality matters
For X to cause Y, the relevant exposure must precede the outcome in the causal process.
This creates particular difficulty for some cross-sectional studies because exposure and outcome are measured at the same point in time. If AI use and writing anxiety are measured simultaneously, for example, the researcher may not know whether AI use preceded anxiety, anxiety preceded AI use, or both developed together.
This does not make cross-sectional studies useless. They can provide valuable descriptive and associational evidence. The problem arises when the design cannot establish the temporal structure required by the particular causal interpretation.
Observational causal inference is not a license for stronger wording
The purpose of causal-inference methodology is not to give researchers permission to replace “association” with “effect.” It is to make causal questions and the assumptions behind their answers more explicit.
A systematic evaluation of causal language in observational health research found substantial ambiguity between ostensibly noncausal wording and causal implications. The authors observed that avoiding explicitly causal words did not necessarily prevent causal interpretation, particularly when recommendations implied that changing the exposure would change the outcome.
This suggests that linguistic caution alone cannot solve an inferential problem. A study can use “associated with” throughout and still make recommendations that implicitly assume causation.
The reverse is also important: observational studies designed explicitly for causal inference should not necessarily be prohibited from stating their causal objective merely because treatment was not randomized. Contemporary methodological proposals instead emphasize making the causal question and assumptions transparent.
“Effect” and “association” should not be treated as interchangeable stylistic choices
If your research question asks:
“What is the effect of academic workload on generative AI use?”
but your analysis estimates only an unadjusted correlation between workload scores and AI-use frequency, there is a mismatch between question and evidence.
If you rewrite the question as:
“Is academic workload associated with generative AI use?”
you have not merely made the sentence more cautious. You have changed the scientific estimand from a causal effect to an observed association.
This distinction follows directly from the issue considered in whether a study can ask “why” without establishing causation. Wording should represent the type of knowledge being sought, not merely disguise an inferential ambition that remains causal underneath.
Prediction is also different from effect
A variable can be highly predictive of an outcome without causing it.
Suppose prior GPA predicts whether students will complete a degree. That does not mean intervening to change the recorded GPA itself would necessarily change degree completion.
Prediction asks whether information about X helps anticipate Y. Causal inference asks what would happen to Y if the relevant exposure or intervention were changed. A predictive model can perform exceptionally well while providing little evidence about causal effects.
Researchers should therefore avoid interpreting predictors automatically as intervention targets.
Some exposures make causal questions difficult to define
Not every characteristic lends itself naturally to an effect question.
Asking for “the effect of age” or “the effect of socioeconomic status” may require considerable conceptual clarification because the exposure is not a simple intervention with one obvious alternative condition. The causal question needs to specify what contrast or intervention is actually being imagined.
This does not mean such questions are impossible. It means the researcher should be explicit about the causal quantity being targeted rather than relying on the word “effect” to carry all the conceptual work.
The research question should reveal whether causal inference is really the objective
Before deciding on terminology, ask what you would do with the answer.
If you want to know whether students who use AI more frequently tend to have different writing scores, association may be exactly the question.
If you want to know whether restricting, providing, or changing AI access would alter writing performance, your substantive question is causal. Avoiding the word “effect” does not make that underlying decision problem noncausal.
Research on causal language has highlighted this tension: studies sometimes use associational terminology while drawing recommendations that only make sense if the relationship is causal.
It is better to formulate the scientific objective explicitly and then determine whether the proposed design can support it.
STROBE does not turn observational research into noncausal research
The STROBE Statement provides reporting guidance for cohort, case-control, and cross-sectional observational studies. Importantly, STROBE describes itself as guidance for reporting observational research rather than a prescription for how studies must be designed or conducted.
Researchers should therefore avoid treating “observational” as a single inferential category. Two cohort studies may use observational data while pursuing quite different objectives: one may estimate an association for prognostic purposes, while another may explicitly attempt to estimate a causal effect.
Your conclusions cannot be stronger than your identification strategy
The decisive issue comes at interpretation.
If your study asks a causal question but the assumptions required for causal interpretation cannot be defended, you may still report the association you successfully estimated. Methodological guidance specifically recommends distinguishing the causal question researchers would like to answer from what their analysis can actually estimate.
This distinction is intellectually useful. A causal question does not become scientifically illegitimate merely because available evidence cannot answer it perfectly. But researchers should not report an associational estimate as though the desired causal effect had been identified.
The question, design, analysis, and conclusion should ultimately agree about what has been learned.