03 · What You Need to Know
The Counterfactual Is the Missing Outcome Behind Every Causal Question
The counterfactual framework is one of the clearest ways to understand what researchers mean by a causal effect. Hernán and Robins introduce causal effects through potential, or counterfactual, outcomes: the outcomes that would occur under different treatment values.
The framework may initially sound abstract. Its logic is quite practical.
Every causal question contains an alternative, even when the wording hides it
Consider the question:
Does formative feedback improve student achievement?
The word “improve” already implies a comparison. Achievement under formative feedback must be compared with achievement under some alternative.
That alternative might be no feedback, usual teaching practice, delayed feedback, another feedback method, or some other clearly defined condition.
Without specifying the alternative, the causal question remains incomplete. “Does it work?” always carries the quieter methodological question: compared with what?
Potential outcomes describe what would happen under each condition
Suppose a student could either receive an intervention or not receive it.
Conceptually, there are two relevant potential outcomes:
- the student's outcome if the student receives the intervention;
- the student's outcome if the student does not receive the intervention.
If the first outcome would be 90 and the second 82, the individual causal effect would be the difference between those potential outcomes: eight points.
The notation can vary across causal-inference traditions, but the underlying idea is the same: causation concerns a contrast between outcomes under alternative conditions.
The fundamental problem is that you cannot usually observe both potential outcomes
Suppose the student actually receives the intervention and scores 90.
You observe the potential outcome under treatment. You do not simultaneously observe what the same student would have scored without treatment.
If the student instead receives no intervention and scores 82, the reverse problem occurs. You observe the no-intervention outcome but not the treatment outcome.
This is sometimes called the fundamental problem of causal inference. For a particular unit at a particular time, researchers generally cannot observe outcomes under mutually exclusive treatment conditions simultaneously.
Factual outcome
The potential outcome corresponding to the condition the unit actually experienced and that can therefore be observed.
Counterfactual outcome
The potential outcome under an alternative condition that the same unit did not actually experience and is therefore unobserved.
The counterfactual is not simply the participant's baseline
Suppose a student's test score is 70 before an intervention and 82 afterward.
It may be tempting to say that 70 represents what the student would have scored without the intervention.
It does not.
The baseline score tells you what was observed earlier. The counterfactual asks what the outcome would have been at the later outcome time under the alternative condition.
The student might have improved from 70 to 76 without the intervention because of ordinary teaching, practice, maturation, or other influences. Or the score might have declined.
This is why a baseline measurement and a counterfactual answer different questions.
A control group is an attempt to provide information about the missing alternative
If you cannot observe the same student simultaneously with and without the intervention, perhaps another group can tell you what would have happened without it.
This is the logic behind many comparison designs.
The crucial issue is whether outcomes in the comparison group provide credible information about the counterfactual outcomes of the intervention group.
If the groups differ systematically in prior achievement, motivation, socioeconomic circumstances, instruction, or other outcome-related characteristics, the comparison group's outcomes may not adequately represent what would have happened to the intervention group under the alternative condition.
This is why merely having a second group is not enough. Researchers must consider whether the comparison group represents an appropriate alternative.
Random assignment creates a particularly powerful counterfactual comparison
In a randomized experiment, treatment assignment is determined by a chance mechanism rather than by participant characteristics or researcher choice.
This makes treatment groups exchangeable with respect to baseline characteristics in expectation. Consequently, the outcomes observed in one randomized condition can provide information about the outcomes that participants in the other condition would have experienced under that alternative assignment.
Researchers still do not observe both potential outcomes for the same individual. Randomization solves the problem at the group level by creating conditions whose outcome distributions can be compared to estimate average causal effects.
This is the deeper reason random assignment strengthens causal inference. Its value is not that it makes every participant identical. It makes the treatment-allocation mechanism suitable for constructing a credible comparison.
Causal research usually estimates average effects rather than individual counterfactuals
Because both potential outcomes cannot generally be observed for one person, researchers commonly focus on average causal effects across a population or study sample.
Conceptually, the average treatment effect compares the average outcome if everyone in the target population received one condition with the average outcome if everyone received the alternative.
Other causal estimands are possible. Researchers may care about effects among those treated, effects in particular subpopulations, time-varying treatments, or other contrasts. The counterfactual logic remains: the effect is defined relative to outcomes under alternative conditions.
Observational causal inference tries to reconstruct the missing comparison under stronger assumptions
Randomization is not always feasible, ethical, or available. Researchers may instead use observational data.
In such studies, treatment or exposure may depend on participant characteristics. People who receive treatment can therefore differ systematically from people who do not.
Causal inference from observational data requires assumptions and methods intended to make observed groups informative about the relevant counterfactual outcomes. Hernán and Robins discuss conditions such as exchangeability, positivity, and consistency as central to identifying causal effects from observational data.
The analytical method alone does not create the counterfactual. Regression adjustment, matching, weighting, standardization, and other methods rely on assumptions about how the observed data can identify the causal contrast.
Exchangeability asks whether one group's outcomes can stand in for another group's missing potential outcomes
Informally, exchangeability means that the groups being compared are sufficiently comparable with respect to the potential outcomes required for the causal contrast.
In a randomized experiment, randomization provides exchangeability in expectation by design.
In an observational study, researchers may seek conditional exchangeability: after accounting appropriately for a sufficient set of pre-exposure confounders, treatment groups are treated as comparable with respect to the relevant potential outcomes.
This is a strong assumption. It cannot generally be verified from the observed data alone because it concerns potential outcomes that are, by definition, partly unobserved.
Positivity asks whether the alternative condition is actually represented
Suppose every participant with a particular characteristic always receives the intervention and no comparable participant ever receives the alternative.
For that part of the population, the data provide no direct information about outcomes under the alternative condition.
Positivity requires, roughly, that relevant units have a nonzero probability of receiving each treatment condition being compared, conditional on the variables needed for the causal analysis.
Severe lack of overlap can make the counterfactual comparison dependent on extrapolation rather than observed comparable cases.
Consistency connects the observed outcome to the relevant potential outcome
When a participant receives the treatment condition being defined, consistency connects that participant's observed outcome with the corresponding potential outcome under that treatment.
This also requires the intervention to be sufficiently well defined for the causal question. “Using AI,” “receiving feedback,” or “participating in tutoring” may describe many substantively different versions of exposure.
If different versions could have different effects, the causal intervention needs to be specified carefully enough for the counterfactual comparison to have a coherent meaning.
Counterfactual reasoning is not the same as imagining any hypothetical scenario
A useful counterfactual must correspond to a meaningful alternative relevant to the research question.
Suppose researchers ask whether attending university causes higher earnings. The alternative “not attending university” still encompasses many possibilities: entering employment immediately, vocational training, unemployment, military service, or another pathway.
Different alternatives can imply different causal contrasts.
The counterfactual should therefore not be treated as a vague imaginary world. It must be tied to a sufficiently defined intervention or exposure contrast.
Counterfactual thinking explains why association alone is insufficient
Suppose tutoring participants earn higher grades than nonparticipants.
That association tells you that outcomes differ between observed groups. A causal claim requires an additional argument: that the nonparticipant outcomes provide valid information about what the tutoring participants would have experienced without tutoring.
If students who choose tutoring are more motivated or academically different, that counterfactual comparison may fail.
This is the distinction underlying association and causation. Association compares what happened to different observed groups. Causal inference asks whether that observed comparison identifies what would have happened under alternative conditions.