03 · What You Need to Know
The Problem Is Causal Meaning, Not Particular Vocabulary
Words carry inferential commitments. If you say that X “affects,” “impacts,” “influences,” or has an “effect on” Y, many readers will reasonably interpret the statement as saying that changing X would produce some change in Y.
That is different from saying that X and Y were associated in the observed data.
A systematic evaluation of observational health research found substantial variation in the causal meaning conveyed by exposure-outcome language. The study also found that recommendations frequently implied stronger causality than the language used to describe the observed relationship. In other words, avoiding an explicitly causal verb did not necessarily prevent researchers from making causal implications elsewhere in the article.
This is why vocabulary rules are an incomplete solution. The question, analysis, interpretation, and recommendations all need to operate at a consistent level of inference.
“Effect” usually carries a clear causal meaning
In causal-inference research, an effect concerns a contrast between outcomes under alternative interventions, exposures, or strategies. A question about “the effect of X on Y” therefore normally asks what would happen to Y if the relevant condition X were changed.
Consider:
“What is the effect of providing students with generative AI access on independent writing performance?”
This is not simply asking whether students with AI access have different writing scores from students without access. It asks whether the alternative condition of providing AI access changes the outcome.
As discussed in the guide on whether a research question can ask about an “effect” without an experimental design, observational data can sometimes support causal-effect estimation. Contemporary methodological guidance explicitly recognizes this possibility but emphasizes the need for a defined causal question, causal estimand, design, assumptions, identification strategy, and careful judgment about whether causal interpretation is tenable.
“Impact” commonly implies that something produced a change
“Impact” is especially common in research titles and questions:
“What is the impact of social media on student mental health?”
“What is the impact of online learning on academic performance?”
“What is the impact of generative AI on critical thinking?”
In ordinary academic usage, these formulations tend to suggest that the exposure produces a consequence. If your study simply measures social media use and mental-health scores in a cross-sectional survey, the word may promise substantially more than the design can establish.
A more appropriate question for an ordinary associational study might be:
“Is social media use associated with depressive symptoms among university students?”
Notice that this is not merely a cautious rewrite. It asks a different question.
“Influence” can be causal even though it sounds softer
Researchers sometimes replace “effect” with “influence” because the latter appears less causal:
“How does academic workload influence students' use of generative AI?”
But “influence” can still imply that workload changes students' behavior. Softening the vocabulary does not necessarily weaken the underlying causal claim.
The same applies to phrases such as “leads to,” “results in,” “contributes to,” “drives,” “increases,” “reduces,” and “shapes.” Their precise causal implication varies by context, but none should be treated as automatically neutral.
Research on causal language illustrates precisely this difficulty: phrases differ in how strongly readers interpret them causally, and causal meaning can also emerge from the surrounding argument rather than from one obviously causal word.
“Associated with” makes a more limited claim
Consider:
“Generative AI use is associated with writing performance.”
This says that values or categories of AI use and writing performance are statistically related in the data under the analysis performed. It does not, by itself, state that changing AI use would change writing performance.
Association
People or cases with different observed values of X also tend to differ in Y under the specified analysis.
Causal effect
The outcome Y would differ under specified alternative conditions of X for a defined target population.
That distinction becomes particularly important when confounding, reverse causation, selection, or measurement processes could explain part or all of an observed relationship.
“Relationship” and “correlation” are not secret synonyms for causation
If the research objective is genuinely associational, other language may also be appropriate:
- Is X associated with Y?
- What is the relationship between X and Y?
- Are X and Y correlated?
- Do levels of Y differ across observed categories of X?
These formulations are not interchangeable in every statistical context, but they generally avoid claiming that manipulating X would alter Y.
“Correlation” has a more specific statistical meaning than the everyday term “relationship,” so use it when the intended analysis and variables make correlation an appropriate description. “Association” is often the more general term.
“Predicts” is not a safer word for “causes”
Another common substitution is “predicts.”
“Does social media use predict academic performance?”
Prediction is a legitimate research objective, but it is different from causal inference. A variable may predict an outcome extremely well without causing it.
For example, a student's prior examination score may predict later examination performance. That does not mean intervening on the recorded prior score would itself improve subsequent performance.
Use predictive language when the purpose is genuinely to predict outcomes and the study evaluates predictive performance appropriately. Do not use “predicts” simply because “affects” feels too strong.
Study design alone does not determine the permissible vocabulary
A widespread rule says:
Experimental study = causal language.
Observational study = associational language.
This is a useful introductory simplification because randomization provides important protection against confounding. It is not a complete account of causal inference.
In 2024, Dahabreh and Bibbins-Domingo proposed a framework for evaluating observational studies that aim to estimate causal effects of interventions. They explicitly argued against deciding whether causal language is permissible solely by checking whether randomization occurred. Instead, they proposed evaluating the causal question, causal estimand, study design, causal assumptions, identification strategy, and whether causal interpretation is ultimately tenable.
This is a more demanding standard than either “observational means no causation” or “we adjusted for covariates, therefore we can say effect.”
Observational causal inference requires an explicitly causal study
Suppose your substantive question is whether restricting smartphone use during class improves student attention. Random assignment may be impossible because some schools already impose restrictions while others do not.
An observational causal study might compare well-defined policy strategies, specify the target population and outcome, establish an appropriate follow-up period, identify potential confounding and selection mechanisms, and use a design and analysis intended to estimate a particular causal effect.
In such a study, causal terminology is not an accidental flourish added during manuscript editing. The causal question organizes the design from the beginning.
Dahabreh and Bibbins-Domingo specifically recommend framing observational causal questions in terms of clearly defined alternative interventions or strategies, outcomes, target populations, and follow-up periods. They also emphasize that causal interpretation remains conditional on assumptions that require substantive evaluation.
Watch Out
Do not decide that an observational study is causal after seeing an interesting adjusted association. Causal inference should shape the question, design, assumptions, data requirements, and analysis rather than being added retrospectively because the result appears persuasive.
Statistical adjustment does not automatically justify “impact” or “effect”
A common argument goes like this: “The study controlled for age, sex, socioeconomic status, and several other variables, so we can now discuss the effect of X.”
That conclusion does not follow automatically.
Adjustment can contribute to causal identification when the selected variables and analytical strategy follow a defensible causal structure. But simply adding available covariates to a regression model does not guarantee that confounding has been adequately addressed. Important confounders may be missing or poorly measured. Adjusting for mediators or colliders can also create problems.
The question is not how many control variables appear in the model. It is whether the assumptions required to interpret the estimate causally are plausible.
Longitudinal does not automatically mean causal either
Longitudinal data can strengthen causal reasoning because the exposure may clearly precede the outcome. That helps with temporality, but temporality is only one requirement.
A longitudinal observational study can still be affected by confounding, selection, measurement error, loss to follow-up, time-varying confounding, and other biases.
Therefore:
“X was measured before Y” does not automatically justify “X influenced Y.”
It may provide a stronger foundation for causal analysis than simultaneous measurement, but the remaining identification problem still matters.
Cross-sectional studies deserve particular caution
In a conventional cross-sectional study, exposure and outcome are often measured at approximately the same time. This can make causal direction especially difficult to establish.
Suppose students with greater writing anxiety report more frequent generative AI use. Did AI use increase anxiety? Did anxious students seek AI assistance more often? Did another characteristic affect both?
If the study cannot distinguish these possibilities, “AI use was associated with writing anxiety” is more defensible than “AI use influenced writing anxiety.”
This is not because the word “influence” is intrinsically forbidden. It is because the study has not established the directional causal claim that the word suggests.
The same word can carry different implications in different sentences
Context matters.
Consider:
“Students perceived that AI influenced how they approached writing.”
Here, “influenced” is explicitly presented as the students' perception. A qualitative interview study could legitimately report that participants described AI in those terms.
Compare:
“AI influenced how students approached writing.”
The qualification has disappeared. The sentence now presents influence as the researcher's substantive conclusion.
This parallels the distinction in asking “why” when a study cannot establish causation. Researchers can study people's causal explanations without automatically endorsing those explanations as established causal effects.
Be careful with “factors affecting” questions
“What factors affect academic performance?” is an extremely common student research question. It is also more demanding than it appears.
The phrase implies that the identified factors produce changes in academic performance. If the study simply administers a questionnaire and examines correlations between several characteristics and grades, it has not necessarily identified “factors affecting” performance.
Depending on the real objective, alternatives might include:
“Which measured characteristics are associated with academic performance?”
or, for a genuine prediction study:
“Which measured characteristics predict subsequent academic performance?”
If the actual goal is causal, however, rewriting the question as an association should not be used to conceal that ambition. The better approach is to acknowledge the causal question and determine whether a suitable design can answer it.
“Impact” is also used in noncausal ways, so context still matters
Academic vocabulary is not perfectly standardized across disciplines. “Impact” may sometimes refer broadly to consequences, significance, societal effects, or participants' perceptions rather than to a formally defined causal estimand.
That ambiguity is precisely why careful wording matters.
If a reader could reasonably interpret “impact of X on Y” as a causal claim, and your study does not support that inference, a more precise term will usually improve the question. There is little methodological benefit in preserving an ambiguous word merely because it is common in the literature.
Do not weaken a genuinely causal question merely to satisfy a vocabulary rule
The opposite mistake is also possible.
Suppose policymakers genuinely need to know whether introducing a school smartphone ban changes classroom attention. The scientific question is causal because the decision concerns what will happen if the policy is implemented.
Writing “Is smartphone-ban status associated with classroom attention?” may describe an observed-data analysis, but it does not fully express the policy question.
Contemporary causal-inference guidance argues that when an observational study genuinely seeks evidence about causal effects, researchers should explicitly state the causal question and assumptions rather than automatically avoiding causal language because treatment was not randomized.
The price of that freedom is methodological accountability. Researchers need to show why the observational comparison can plausibly inform the causal question.
Recommendations can reveal hidden causal claims
Imagine a paper concluding:
“Social media use was associated with lower academic performance. Universities should therefore reduce students' social media use to improve grades.”
The first sentence is associational. The recommendation is causal. It assumes that intervening to reduce social media use would improve grades.
The systematic evaluation by Haber and colleagues found precisely this kind of mismatch: action recommendations often implied stronger causality than the language linking exposures and outcomes.
Researchers should therefore audit more than the research question. The abstract, discussion, conclusion, practical implications, and recommendations should all respect the same inferential boundaries.
Reporting guidelines do not replace causal reasoning
STROBE provides reporting recommendations for cohort, case-control, and cross-sectional observational studies. The STROBE initiative explicitly states that its recommendations concern how observational studies should be reported and are not prescriptions for designing or conducting them.
Following STROBE can improve transparency, but it does not determine whether a particular estimate is causal. That judgment requires examining the research question, design, assumptions, measurement, analysis, and possible sources of bias.
A useful vocabulary depends on what you are actually asking
| Research purpose |
Wording that may fit |
What the study must support |
| Description |
What is the prevalence, frequency, level, or distribution of X? |
A defensible description of the defined population or cases |
| Association |
Is X associated with Y? What is the relationship between X and Y? |
An appropriately estimated observed relationship |
| Comparison |
Does Y differ between observed groups A and B? |
A defensible comparison without automatically attributing the difference to group membership |
| Prediction |
Does X predict Y? Which characteristics predict Y? |
Appropriately evaluated predictive performance |
| Causal effect |
What is the effect of X on Y? Does X increase or reduce Y? |
A defined causal contrast and an identification strategy capable of supporting causal interpretation |
| Participants' perceived influence |
How do participants perceive X as influencing Y? |
Evidence about participants' interpretations and experiences, not automatic proof of the causal relationship |
The table is not a dictionary of permitted and prohibited words. It is a reminder that wording should correspond to the inferential task.
Consistency matters more than linguistic caution alone
You can write an entirely associational research question and then overclaim causality in the conclusion. You can also formulate an explicitly causal observational question and conduct a carefully designed causal analysis.
What matters is alignment.
The question should say what you want to know. The design should generate evidence appropriate to that question. The analysis should estimate the relevant quantity. The assumptions should be explicit enough to evaluate. The conclusion should not exceed what those elements jointly support.
Once those pieces align, choosing between “association,” “effect,” “influence,” or another term becomes much less mysterious.