03 · What You Need to Know
The Difference Is Not the Statistical Significance of the Result
Association and causation are not two levels on a p-value scale.
A very small p-value can accompany a non-causal association. A modest causal effect can have substantial statistical uncertainty. The inferential distinction comes from the question, design, assumptions, and interpretation rather than from whether a coefficient crosses a conventional significance threshold.
Association asks whether observed variables differ or vary together
An associational question might ask whether students who use a tutoring system more often tend to earn higher grades.
The researcher can quantify that relationship using methods appropriate to the variables and design. Depending on the study, this might involve correlation, regression, contingency tables, group comparisons, or other models.
If the relationship is estimated appropriately, the study may provide useful evidence even without causal interpretation.
Association is not a consolation prize. Many important scientific questions are genuinely descriptive or associational.
Causation asks what would happen under an alternative condition
A causal question changes the target:
Would students' grades change if their tutoring exposure were changed?
Now the researcher needs a comparison between potential outcomes under alternative exposure conditions.
This is why understanding the distinction among correlation, association, prediction, and causation matters before interpreting the analysis. A relationship observed in the data and an effect of intervening on a variable are different quantities.
Ask how the exposure was determined
One of the fastest ways to diagnose a causal problem is to ask why some participants were exposed and others were not.
If exposure was randomly assigned, baseline characteristics do not systematically determine treatment allocation under the randomization mechanism.
If exposure was self-selected, chosen by clinicians, determined by teachers, influenced by institutional resources, or otherwise naturally occurring, group membership may encode important differences that also affect the outcome.
That does not automatically rule out causal inference. It tells you that the causal analysis must address the nonrandom treatment-assignment process explicitly.
If important common causes remain uncontrolled, the observed association may be confounded
Suppose motivation affects both voluntary tutoring use and academic performance.
Students with greater motivation may use tutoring more frequently and study more effectively outside the platform. A positive tutoring-achievement association could therefore partly reflect motivation rather than a causal effect of tutoring.
Adding measured confounders to a model can help when those variables are correctly identified, measured, and modeled. But a causal conclusion requires an argument that the relevant confounding has been addressed sufficiently for the intended estimand.
If a major common cause was never measured, the design may support association more comfortably than causation.
Unclear temporal order is a major warning sign
If X is proposed to cause Y, the relevant X must precede Y.
Suppose a cross-sectional survey finds that students reporting greater use of generative AI also report lower confidence in academic writing. Did AI use reduce confidence, or did students with low confidence turn to AI more often?
If the timing is unresolved, both directions may fit the observed association.
Without adequate temporal ordering, the study may not distinguish the intended causal explanation from reverse causation.
A before-and-after design can show change without identifying its cause
Suppose anxiety scores decline after students complete a stress-management workshop.
The study establishes that measured anxiety changed over the period among those participants.
But what would their anxiety have been at follow-up without the workshop?
Perhaps it would also have declined because examinations ended, workloads decreased, participants became familiar with the questionnaire, or other events occurred.
A baseline provides the starting value. It does not by itself provide the missing counterfactual outcome.
Having a comparison group does not automatically establish causation
Suppose the intervention group comes from one university and the comparison group from another.
The universities differ in admissions, teaching practices, student demographics, resources, and assessment policies. A difference in outcomes may reflect any combination of those factors.
The mere existence of two groups therefore does not establish a causal comparison.
The question is whether the comparison group provides credible information about the relevant alternative outcome.
Baseline similarity does not prove exchangeability
Researchers sometimes show that exposed and unexposed groups have similar means on several measured variables and conclude that confounding has been eliminated.
That conclusion is too strong.
Measured similarity is informative, but unmeasured differences may remain. The absence of statistically significant baseline differences also does not demonstrate equivalence.
In nonrandomized research, exchangeability is an identifying assumption about potential outcomes, not a property that can be certified from a short demographic table.
Adjustment does not automatically convert association into effect
A multiple regression may produce an “adjusted association.” Whether that coefficient has a causal interpretation depends on why those covariates were included and what causal structure is assumed.
Adjusting for a mediator can block part of the causal pathway. Adjusting for a collider can induce an association that was not present in the relevant causal contrast. Adjusting for variables measured after exposure can create other complications.
The number of covariates in the model is therefore not a causal-validity score.
Prediction can be excellent while causal interpretation is poor
A model might predict student dropout extremely well using variables that are not causes of dropout.
For example, a variable may be a proxy for an underlying risk process or may occur downstream of earlier causes. It can still contribute substantial predictive information.
If the goal is identifying students at risk, that may be perfectly useful.
If the goal is deciding what to intervene on, predictive importance alone is insufficient. Changing a strong predictor does not necessarily change the outcome.
Cross-sectional designs often support association more readily than causal interpretation
Cross-sectional data commonly measure exposure and outcome at approximately the same time. When the temporal histories of the variables are unknown, reverse causation becomes difficult to exclude.
Confounding and selection may also remain.
This does not mean every cross-sectional variable lacks temporal information or that causal inference is logically impossible from every cross-sectional dataset. It means the design often leaves key causal requirements unresolved, making associational language more defensible.
Longitudinal data solve only part of the problem
Measuring exposure before outcome can satisfy an important temporal requirement.
It does not ensure that exposed and unexposed participants are comparable.
A longitudinal observational study in which treatment is self-selected can still suffer from confounding. Attrition can create additional selection problems, and time-varying confounders can complicate conventional adjustment.
“Longitudinal” therefore describes temporal structure, not causal identification.
Some observational designs are explicitly constructed for causal inference
It would be equally mistaken to conclude that every observational study must stop at association.
Researchers may design observational analyses around explicit causal estimands, target trial emulation, natural experiments, instrumental variables, regression discontinuity, difference-in-differences, or other strategies when their assumptions are appropriate.
The relevant question is when observational evidence can support a causal claim, not whether observational data carry a permanent “association only” label.
A useful diagnostic is to ask what makes the comparator represent the missing counterfactual
Suppose users of an educational application perform better than non-users.
Ask:
Why should the outcomes of the non-users tell me what would have happened to the users if the users had not used the application?
If the answer is simply “because both groups are in the dataset,” the causal argument is weak.
If the answer involves randomized assignment, or a carefully justified observational identification strategy with explicit assumptions, then a causal interpretation may be more defensible.
The language in your paper should reveal the strength of your claim
Researchers can unintentionally shift from association to causation through verbs.
| If the evidence establishes... |
Language that may fit |
Language requiring causal support |
| Statistical relationship |
was associated with, was related to, differed between, correlated with |
caused, produced, led to |
| Prediction |
predicted, was predictive of, improved predictive performance |
prevented, reduced, increased |
| Temporal association |
preceded, was prospectively associated with |
resulted in, produced |
| Defensible causal effect |
increased, reduced, affected, caused, had an effect on |
The strength and specificity should still match the estimand and evidence. |
There is nothing inherently weak about precise associational language. The problem arises when causal verbs imply an intervention effect that the design never identified.
The abstract and conclusion are common places for causal overreach
A paper may use careful associational terminology throughout the results and then conclude that an exposure “improves,” “reduces,” or “leads to” an outcome.
The statistical model did not become more causal between the results and conclusion sections.
Your conclusion should preserve the inferential level established by the design. If the study estimates association, discuss the association, its plausible explanations, and what further evidence would be needed to evaluate causality.
Reporting guidelines improve transparency but do not decide causality for you
STROBE provides reporting guidance for major observational designs, including cohort, case-control, and cross-sectional studies. It encourages transparent reporting of what was planned, done, found, and interpreted.
STROBE also explicitly notes that its recommendations are not prescriptions for designing or conducting studies and that the checklist is not an instrument for evaluating research quality.
Using the correct reporting guideline is important. It does not determine whether your coefficient should be called a causal effect.