01 · The Question
What Are You Actually Trying to Establish?
Suppose you find that students who use a particular learning platform more frequently also tend to earn higher course grades. What have you established?
You might say the two variables are correlated. You might describe platform use as associated with achievement. Perhaps you could use platform activity to predict which students are likely to perform well. But can you conclude that using the platform improves achievement?
That last step is where the research question changes.
Correlation, association, prediction, and causation are related ideas, but they do not make interchangeable claims. A study can demonstrate a statistical relationship without explaining why it exists. A variable can be useful for predicting an outcome without causing that outcome. And a causal question requires more than finding a statistically significant coefficient in a regression model.
The important question is therefore not simply which statistical test you plan to run. It is what kind of claim your research is designed to support.
03 · What You Need to Know
Four Similar-Sounding Ideas That Answer Different Questions
The easiest way to distinguish these concepts is to stop thinking first about statistical procedures and instead ask what you want to know about the variables.
Correlation asks whether two quantitative variables vary together
In its common statistical sense, correlation summarizes the direction and strength of a relationship between variables. For example, a researcher might examine whether the number of hours students spend studying is correlated with their examination scores.
A positive correlation means that higher values of one variable tend to occur with higher values of the other. A negative correlation means that higher values of one tend to occur with lower values of the other. A correlation near zero indicates little or no relationship of the particular form measured by the chosen correlation coefficient.
Correlation is therefore a description of how variables co-vary. It does not, by itself, explain the mechanism producing that pattern.
It is also useful to remember that correlation is narrower than association. Common correlation coefficients, such as Pearson's correlation coefficient, describe particular forms of statistical relationship. Two variables may have an association that is not adequately represented by a particular correlation coefficient.
Association is the broader idea of variables being statistically related
Association is a more general term. If the distribution or probability of one variable differs depending on another variable, researchers may describe them as associated.
For instance, employment status might be associated with participation in an adult-learning program. Because both variables could be categorical, describing their relationship as a correlation may not be the most natural terminology. Association accommodates a wider range of variables and analytical approaches.
Correlation
A particular way of describing how variables vary together, commonly through a correlation coefficient.
Association
The broader concept that variables are statistically related or dependent in some way.
Neither term automatically tells you why the relationship exists. An observed association might reflect a causal effect, reverse causation, confounding, selection processes, measurement problems, chance, or some combination of these possibilities.
Prediction asks whether available information can forecast an outcome
Prediction has a different objective. A predictive study asks whether information available about a case can be used to estimate an outcome for that case, ideally with useful performance when applied beyond the data used to develop the model.
Imagine that a university wants to identify students at elevated risk of dropping out. A model might combine attendance, previous grades, learning-management-system activity, financial information, and other variables to estimate that risk.
The predictors do not all have to cause dropout to be useful predictors. A variable may contain information about the outcome because it is related to an underlying cause, reflects an earlier part of the same process, or acts as a proxy for something else.
This distinction matters because prediction and causal explanation optimize for different goals. Research on the conflation of prediction and causal inference has shown that researchers can make methodological errors when variables chosen for predictive usefulness are interpreted causally, or when causal analyses select variables merely because they predict the outcome well.
Causation asks what would happen if something changed
A causal question goes beyond observing that X and Y occur together. It asks whether Y would differ if X were changed, compared with what would have happened under an alternative condition.
Consider two questions:
- Are students who receive more formative feedback more likely to achieve higher final scores?
- Would providing students with more formative feedback improve their final scores?
The first can be framed as an associational question. The second is causal. It asks about the consequence of changing an exposure or intervention.
This is why causal inference is closely connected to the idea of a counterfactual. For a student who received additional feedback, we can observe the outcome under that condition, but we cannot simultaneously observe what the same student's outcome would have been at the same time had the student not received it. Causal research attempts to construct a defensible comparison that represents this unobserved alternative.
The same variables can support very different research questions
Suppose a dataset contains weekly study time and final examination scores. Researchers could ask several questions using exactly those variables.
| Research aim |
Example question |
What the result is intended to establish |
| Correlation |
How strongly are weekly study hours correlated with examination scores? |
The direction and strength of co-variation measured by the selected correlation coefficient |
| Association |
Are study habits associated with examination performance? |
Whether the variables are statistically related |
| Prediction |
Can study behavior help predict a student's examination score? |
Whether available information can forecast the outcome with useful predictive performance |
| Causation |
Would increasing students' study time improve their examination scores? |
The expected change in the outcome under a change in the exposure |
The dataset alone does not determine which question you are answering. Your research objective, design, assumptions, and analytical strategy do.
Statistical adjustment does not automatically turn association into causation
A common source of confusion is regression analysis. Researchers sometimes assume that an independent variable becomes a cause once other variables have been entered as covariates.
That is not what regression itself establishes.
Regression can be used for descriptive, associational, predictive, and causal analyses. What changes is the purpose of the model, how variables are selected, the assumptions being made, and how the resulting estimates are interpreted.
For causal inference, confounding is particularly important. A confounder is a factor related to both the exposure and outcome in a way that can distort the exposure-outcome relationship. Deciding what should be adjusted for therefore requires substantive knowledge and causal reasoning, not simply asking software to identify statistically significant covariates.
Nor is adjusting for every available variable necessarily safer. Depending on the causal structure, controlling for variables such as mediators or colliders can alter or bias the estimate researchers are trying to interpret causally.
Temporal order matters, but it is not sufficient
A proposed cause must precede its effect. If an exposure is measured after the outcome has already occurred, the intended causal interpretation may become difficult or impossible to defend.
Establishing which variable came first, however, does not eliminate confounding or other explanations. Longitudinal data may strengthen temporal reasoning compared with a purely cross-sectional snapshot, but longitudinal does not mean causal by definition.
Experimental designs can strengthen causal inference, but the word experimental is not magic
Randomized experiments are powerful for causal questions because random assignment can make treatment groups comparable on both measured and unmeasured baseline characteristics on average, subject to chance variation. This helps separate the intervention from competing explanations for differences in outcomes.
Even then, researchers still need to consider issues such as attrition, noncompliance, measurement, treatment implementation, interference between participants, missing data, and whether the findings apply to the population or setting of interest.
Random assignment should also not be confused with random sampling. The former concerns assignment to conditions and can strengthen internal causal inference. The latter concerns how units are selected from a population and is more directly related to sampling and generalizability.
Observational does not necessarily mean causation is forever off-limits
It is equally misleading to adopt the opposite rule and declare that observational research can never contribute to causal inference.
Modern causal inference includes methods for estimating causal effects from observational data under explicit assumptions. The difficulty is that those assumptions, particularly those concerning confounding, may be demanding and cannot generally be verified from the observed data alone.
Whether an observational study can support a causal claim therefore depends on considerably more than the label attached to the study design. Researchers need to articulate the causal question, define the relevant comparison, establish temporal ordering, identify plausible confounders, choose an appropriate analytical strategy, and examine how sensitive the conclusion may be to violations of key assumptions.
04 · A Practical Example
One Dataset, Four Very Different Claims
Hypothetical Example
Does participation in an AI tutoring system improve mathematics performance?
A researcher has data from 2,000 university students. The dataset includes the number of AI tutoring sessions each student completed during the semester, prior mathematics achievement, demographic variables, course attendance, and final examination scores.
Correlation The researcher calculates a correlation between the number of tutoring sessions and final examination scores. Students who use the system more frequently tend to have higher scores. This establishes a pattern of co-variation, not why the pattern exists.
Association The researcher fits a statistical model and finds that tutoring-system use remains associated with examination performance after accounting for several measured variables. The adjusted relationship may be informative, but adjustment alone does not establish that tutoring caused the difference.
Prediction The researcher develops and validates a model using tutoring activity together with other student information to predict final examination performance. If the model predicts well in appropriate new data, tutoring activity may be a useful predictor even if it is not itself a cause of higher performance.
Causation The researcher instead asks what students' examination performance would have been if they had received or used the tutoring intervention compared with an appropriate alternative. Answering this requires a design and analysis capable of supporting that counterfactual comparison and addressing plausible competing explanations.
Why might the original positive relationship be misleading as evidence of a tutoring effect? Students who voluntarily use the system more frequently may already be more motivated. Perhaps they attend more classes, have stronger prior knowledge, or seek help more actively. Alternatively, struggling students might use the system more often, which could push the observed association in the opposite direction.
These possibilities are not statistical trivia. They represent different explanations for the same observed relationship.
If the researcher wants to estimate the causal effect of providing access to the tutoring system, a randomized study might be appropriate when feasible and ethical. If randomization is unavailable, an observational causal analysis would require a clearly specified causal question and defensible assumptions about how treatment selection and confounding are handled.
06 · What This Means for You
Choose the Claim Before You Choose the Analysis
One of the most useful decisions you can make early in a study is to write down the sentence you ultimately hope your evidence will allow you to say.
If that sentence is essentially “X and Y are related,” you are asking an associational question. If it says “we can estimate Y using X,” your goal is predictive. If it says “changing X would change Y,” you have entered causal territory.
Your design should follow from that distinction rather than discovering, after the analysis, that a convenient statistical result has become a much stronger claim than the study was built to support.
A simple decision framework
If you want to describe how two quantitative variables move together
Frame the question around correlation and select a correlation measure appropriate to the variables and form of relationship.
If you want to know whether variables are statistically related
Frame the study around association and choose a design and analysis suited to the variables and population.
If you want to forecast an outcome for new cases
Treat prediction as the primary goal and evaluate predictive performance using appropriate validation rather than interpreting every predictor as a cause.
If you want to know what would happen if an exposure, treatment, policy, or condition changed
State the causal question explicitly and design the study around the relevant counterfactual comparison, temporal order, confounding structure, assumptions, and causal estimand.
Watch Out
Do not upgrade the language of your conclusion after seeing an interesting result. If the research question, design, and analysis were built to establish association, a strong or statistically significant association does not by itself license verbs such as “caused,” “improved,” “reduced,” “increased,” or “led to.”
If your causal question requires comparing outcomes under different conditions, you may also need to consider whether the research question actually requires a comparison and, if so, what comparison would make the causal contrast meaningful. Those are design decisions, not details to be added after data collection.
07 · A Quick Checklist
Before You Describe Your Study as Correlational, Predictive, or Causal
Before finalizing your research question and design, check:
Can you state clearly whether your primary aim is to describe a relationship, predict an outcome, or estimate the effect of changing something?
Does the wording of your research question match the strength of claim your design can reasonably support?
If your aim is prediction, have you planned an appropriate way to evaluate performance beyond merely fitting the model to the development data?
If your aim is causal, have you defined the exposure or intervention, outcome, target population, and relevant alternative condition clearly?
For a causal question, does the proposed cause occur before the outcome?
Have you considered plausible confounders and alternative explanations rather than relying only on variables selected by statistical significance?
Does your analytical method serve your research aim rather than determine the aim retrospectively?
Will the language in your abstract, results, discussion, and conclusion remain consistent with what the design actually establishes?