03 · What You Need to Know
Association Does Not Tell You Which Way the Arrow Points
Begin by separating association from direction
Suppose X and Y are positively correlated.
At minimum, the data indicate that values of X and Y vary together under the conditions studied.
Several explanations may produce that pattern:
X → Y
Y → X
X ↔ Y
or:
Z → X and Z → Y
These structures can all generate an X–Y association.
The observed association therefore does not contain an automatic arrow.
Association
X and Y vary together.
Directionality
Theoretical or empirical evidence specifies whether X precedes or contributes to Y, Y precedes or contributes to X, or both processes occur.
Theory can propose direction even when the current data cannot prove it
Researchers are not required to abandon all directional hypotheses simply because the study is observational or cross-sectional.
A well-developed theory may clearly propose:
X → Y
The conceptual framework can represent that proposition as a hypothesis.
The crucial distinction is between:
“Our theory proposes that X contributes to Y.”
and:
“Our study demonstrates that X causes Y.”
The first describes the model being evaluated. The second describes the strength of inference supported by the design.
Those statements need not be identical.
Known temporal facts can rule out some directions
Sometimes direction is clear because one variable necessarily precedes the other.
Date of birth precedes current university enrollment. A randomly assigned intervention occurs before post-intervention outcomes. Prior academic records exist before an outcome measured years later.
Such timing can make the reverse pathway impossible or implausible for the particular causal contrast.
But temporal order alone is not sufficient for causation.
If X occurred before Y, X may still be unrelated causally to Y, or a third variable may have caused both.
Measurement order is not the same as causal order
This distinction is easy to miss.
Suppose intelligence is measured at Time 2 and school achievement was measured at Time 1. It would be unreasonable to conclude that achievement caused intelligence merely because achievement was measured first.
The underlying constructs existed before their measurement occasions.
Similarly, measuring motivation in September and engagement in October does not automatically prove that September motivation caused October engagement. Prior engagement may already have shaped September motivation.
Watch Out
“X was measured first” does not necessarily mean “X occurred first in the causal process.” Temporal design should reflect the timing of the underlying phenomenon, not merely the order in which questionnaires were administered.
Cross-sectional data often provide weak evidence about temporal ordering
When X and Y are measured simultaneously, establishing which one preceded the other is often difficult.
For example, suppose self-efficacy and engagement are strongly correlated in a one-time student survey.
The data may be compatible with:
Self-efficacy → Engagement
because confidence encourages effort and persistence.
They may also be compatible with:
Engagement → Self-efficacy
because successful participation creates mastery experiences.
Or both processes may operate.
Unless the design or external knowledge provides additional information, the cross-sectional association alone generally cannot settle the direction.
Regression does not establish direction
Researchers sometimes fit:
Y = b₀ + b₁X + e
and interpret X as influencing Y because X appears on the right side of the equation.
But statistical placement is not causal evidence.
In many ordinary cross-sectional settings, reversing the variables and modeling X as a function of Y does not resolve which causal process generated their association.
The designation “independent variable” can therefore become misleading when it is interpreted as “variable proven to cause the dependent variable.”
Structural equation modeling does not automatically solve directionality either
Structural equation models can represent directional hypotheses and compare complex theoretical structures. This is useful.
But drawing X → Y in SEM does not independently establish that X causes Y.
Alternative directional models may sometimes fit similarly. Model fit indicates compatibility between the specified model and observed covariance structure, subject to assumptions. It does not guarantee that the arrows reproduce the real causal process.
Direction still requires substantive and design-based justification.
A better-fitting directional model is evidence, not final proof
Suppose researchers compare:
Model A: X → Y
and:
Model B: Y → X
and Model A fits better.
That result may contribute evidence under the model assumptions, but it does not automatically eliminate every alternative explanation. Unmeasured confounding, measurement error, model misspecification, and reciprocal processes may remain.
Model comparisons should therefore be interpreted as part of an argument about direction rather than as a machine that discovers causality from covariance alone.
Longitudinal data can strengthen directional reasoning
Repeated measurements make it possible to examine whether earlier variation in one construct predicts later variation in another while accounting, in some models, for prior levels of those constructs.
For example, researchers might measure self-efficacy and engagement at several time points and examine whether:
Earlier self-efficacy predicts later engagement
and whether:
Earlier engagement predicts later self-efficacy.
This provides substantially more temporal information than measuring both variables once. Longitudinal data can help distinguish plausible unidirectional, reciprocal, or more complex relationships.
However, longitudinal measurement does not automatically establish causal effects. The model must still address confounding, measurement, selection, and appropriate separation of temporal processes.
Cross-lagged models require careful interpretation
Cross-lagged panel models have long been used to examine reciprocal longitudinal relationships. A conventional model may estimate X at Time 1 predicting Y at Time 2 while accounting for earlier Y, alongside Y at Time 1 predicting X at Time 2 while accounting for earlier X.
This can be informative, but contemporary longitudinal methodology emphasizes that between-person differences and within-person changes should not be conflated.
For example, people who generally have high self-efficacy may also generally have high engagement. That between-person association is not the same as asking whether a temporary increase in one person's self-efficacy predicts a later change in that same person's engagement.
The statistical model should match the level of causal or developmental process the theory actually proposes.
A longitudinal lag should correspond to the process
Suppose X is theorized to affect Y within hours, but researchers measure the variables once per year.
Alternatively, suppose the process unfolds over several years but measurements are taken one week apart.
Neither schedule may capture the relevant causal dynamics adequately.
Longitudinal research is not strengthened simply by adding “Time 1” and “Time 2” labels. The interval between measurements should be substantively plausible for the process under study.
Experiments can clarify direction when the focal variable can be manipulated
If researchers can manipulate X and then measure subsequent Y under a well-designed randomized experiment, the direction from intervention to outcome becomes much more defensible.
For example, random assignment to a feedback intervention occurs before later self-efficacy and achievement outcomes.
Because assignment is controlled by the study, reverse causation from the later outcome to treatment assignment is not a plausible explanation.
Other issues can remain, including noncompliance, attrition, measurement problems, mediation assumptions, and questions about generalizability, but randomization provides a much stronger basis for causal direction than an ordinary cross-sectional association.
Not every variable can or should be manipulated
Some proposed causes are not manipulable in a straightforward sense: age, historical exposure, ethnicity, geographic origin, or many social conditions.
Causal questions about such variables require careful definition and appropriate observational or quasi-experimental reasoning rather than the simplistic rule that only randomized variables can have causal consequences.
The important point is that the design should provide evidence relevant to the causal contrast being claimed.
Reverse causality deserves explicit consideration
Reverse causality occurs when a relationship interpreted as X → Y may actually arise partly or wholly because Y influences X.
Suppose researchers find that faculty who use more educational technology report higher technology self-efficacy.
It is tempting to conclude:
Self-efficacy → Technology use.
That is theoretically plausible.
But repeated technology use may also create mastery experiences that increase self-efficacy:
Technology use → Self-efficacy.
A strong discussion should consider this possibility rather than treating the selected arrow as self-evident.
Reverse causality is not the same as confounding
Suppose X and Y are associated.
Reverse causality proposes:
Y → X
Confounding proposes something more like:
Z → X
Z → Y
Both can threaten a simple causal interpretation of X → Y, but they are different explanations.
A study may need to consider both.
Reciprocal causation may be theoretically more realistic
Some relationships are not well represented by choosing one winner in a directional contest.
Consider:
Engagement → Achievement
Students who engage more may learn more.
But:
Achievement → Engagement
Successful performance can also motivate further engagement.
Across time, both pathways may operate.
If theory predicts this feedback, researchers should consider whether the relationship is genuinely reciprocal rather than simply unclear.
Do not interpret reciprocal possibility as permission to draw two arrows casually
A bidirectional arrow can become another way of avoiding theoretical commitment.
If both directions are proposed, explain why each pathway is plausible and what temporal process connects them.
For example:
Higher self-efficacy encourages engagement, while successful engagement experiences subsequently strengthen self-efficacy.
This describes a feedback mechanism.
Simply writing “X and Y influence each other” without specifying when or how offers little explanatory value.
Mediation requires particularly careful directionality
A mediation model proposes:
X → M → Y
If M and Y could plausibly be reversed, the meaning of the mediation claim changes fundamentally.
For example:
Stress → Self-efficacy → Engagement
and:
Stress → Engagement → Self-efficacy
represent different mechanisms.
Estimating one indirect effect does not demonstrate that its ordering is correct. This is why claims that a mediator explains a relationship require attention to temporal and causal ordering.
Antecedent terminology should also reflect actual ordering
Calling X an antecedent of Y implies that X comes earlier in the relevant theoretical or temporal sequence.
If the study cannot justify that ordering, the label may overstate what is known.
This is why understanding what counts as an antecedent variable requires more than observing a significant regression coefficient.
Causal diagrams can expose competing directional assumptions
Drawing alternative causal structures can clarify exactly what is uncertain.
For example:
Model 1: X → Y
Model 2: Y → X
Model 3: X ↔ Y over time
Model 4: Z → X and Z → Y
The exercise forces researchers to ask what evidence would discriminate among these possibilities.
A causal diagram does not solve the uncertainty by itself, but it turns a vague concern about “direction” into specific competing explanations.
Sometimes your present study cannot resolve direction
This is an important scientific conclusion, not a methodological embarrassment.
If theory supports several directions and the current cross-sectional design cannot distinguish among them, the appropriate response may be:
The observed association is consistent with the proposed X → Y relationship, but reverse or reciprocal relationships cannot be excluded.
That conclusion is more defensible than pretending that a directional regression coefficient answered a question the design could not settle.
Use language that matches the remaining uncertainty
If direction is unresolved, phrases such as:
- “was associated with”;
- “was related to”;
- “covaried with”;
- “was a statistical predictor of” when prediction is genuinely intended;
may be more appropriate than:
- “caused”;
- “affected”;
- “led to”;
- “resulted in.”
The distinction among association, influence, effect, and prediction therefore becomes particularly important when directionality is uncertain.
Your conceptual framework may still show a directional hypothesis
There is a difference between proposing a theory and claiming the theory has already been proven.
If theory clearly predicts X → Y, a conceptual framework may represent that direction even when the current design provides only limited causal evidence.
The manuscript should then be explicit:
The framework hypothesizes X → Y, while the study tests relationships consistent with that model rather than definitively establishing causal direction.
This distinction keeps theoretical clarity without exaggerating empirical certainty.
If no direction is theoretically privileged, a nondirectional formulation may be better
Sometimes theory and evidence establish only that X and Y should be related.
In such cases, forcing a directional hypothesis may add assumptions without adding understanding.
A nondirectional research question such as:
What is the relationship between X and Y?
may be more appropriate.
Direction can then become a question for future longitudinal, experimental, or otherwise causally informative research.
Directional uncertainty can improve the next study
Rather than treating unresolved direction as a limitation sentence added at the end of a paper, researchers can use it to design the next investigation.
Possible improvements include:
- measuring variables repeatedly;
- using theoretically meaningful time lags;
- including prior levels of the relevant constructs;
- randomizing an intervention where appropriate;
- using quasi-experimental variation;
- collecting measures of plausible common causes;
- testing competing directional models transparently.
Uncertainty can therefore become a design question rather than merely a disclaimer.