03 · What You Need to Know
“Observational” Describes How Exposure Occurred, Not Necessarily the Question You Are Allowed to Ask
In an observational study, researchers observe exposures or treatments that were not assigned through the study's randomization procedure. That distinction matters because the exposed and unexposed groups may differ systematically before the outcome occurs.
Yet observational research can still be organized around a causal estimand. Contemporary causal-inference frameworks explicitly use observational data to estimate causal effects when randomized trials are unavailable, infeasible, unethical, or otherwise unsuitable.
The difficult part is establishing why the observed comparisons identify the causal effect of interest.
Begin with a causal question, not with an available regression coefficient
A causal analysis should start by defining what effect you actually want to estimate.
Compare:
- Are students who frequently use an AI tutoring system more likely to pass mathematics?
- What would happen to mathematics pass rates if eligible students used the AI tutoring system regularly compared with if they did not use it?
The first asks about an observed association. The second defines a causal contrast between alternative exposure conditions.
The distinction matters because the second question requires researchers to reason about counterfactual outcomes. You need evidence about what outcomes would have occurred under exposure conditions different from those actually observed.
Specify the causal estimand before deciding how to estimate it
A causal effect is not simply “the effect of X on Y.” Researchers need to define the population, treatment or exposure strategies, outcome, relevant time horizon, and causal contrast.
Recent methodological frameworks for observational causal inference emphasize this explicit specification. One such framework organizes the problem around the target estimand, target population, target trial, and target validity.
That order is useful because an analytical method cannot tell you what scientific effect you intended to estimate. The estimand must come first.
The target trial provides a useful design benchmark
One influential approach asks researchers to describe the hypothetical randomized trial that would answer the causal question if such a trial were feasible.
This hypothetical experiment is called the target trial.
Researchers specify features such as:
- eligibility criteria;
- treatment or exposure strategies;
- assignment procedures;
- the start of follow-up;
- outcomes;
- the causal contrast of interest;
- the analysis plan.
The observational study is then designed to emulate these components as closely as the available data permit.
Target trial emulation does not turn observational data into randomized data. Its value is conceptual and methodological: it makes the causal question explicit and can help researchers avoid design errors that arise when eligibility, treatment assignment, baseline, and follow-up are defined inconsistently.
Temporal order must support the proposed causal direction
A proposed exposure must occur before the outcome it is claimed to affect.
This can be difficult when observational data were collected without the causal question in mind. Exposure and outcome may be recorded concurrently, or an apparent baseline may occur after exposure has already begun.
Researchers should therefore establish the relevant temporal order rather than assuming that a longitudinal dataset automatically provides it.
The start of follow-up is particularly important. If eligibility, treatment classification, and follow-up begin at inconsistent times, researchers can introduce avoidable biases even when the subsequent statistical model is sophisticated.
Confounding is one of the central challenges
Without random assignment, exposure may be related to characteristics that also affect the outcome.
Suppose students who voluntarily use an academic-support platform are more motivated, have stronger prior achievement, or receive more instructor encouragement. Those characteristics could influence both platform use and later achievement.
If they are not appropriately addressed, the observed exposure-outcome difference may not represent the causal effect of the platform.
This is confounding.
The challenge is not simply to “control for some variables.” Researchers need to identify a sufficient set of pre-exposure confounders using substantive and causal knowledge and then measure and incorporate them appropriately.
More covariate adjustment is not necessarily better causal adjustment
A familiar observational analysis includes every variable available in the dataset and interprets the adjusted exposure coefficient as causal.
That strategy is not generally justified.
Variables can occupy different positions in the causal structure. Some are confounders. Others may be mediators caused by the exposure. Some can be colliders influenced by variables in the causal system. Adjusting indiscriminately can change the causal effect being estimated or introduce bias.
Variable selection should therefore follow the causal question and assumed causal structure rather than a rule such as “include everything associated with the outcome” or “adjust for everything significant in the univariate analysis.”
No unmeasured confounding is a strong assumption
Even careful adjustment can address only the relevant variables that were measured adequately.
For many observational causal analyses, researchers rely on an exchangeability assumption that, conditional on an appropriate set of measured variables, the exposure groups are sufficiently comparable with respect to their potential outcomes.
This assumption is demanding because important confounders may be unavailable, poorly measured, or unknown.
Crucially, a good balance table after matching or weighting cannot prove that no unmeasured confounding remains. The table describes measured information. The causal assumption extends beyond what is directly observed.
Positivity asks whether meaningful comparisons actually exist in the data
Suppose every student with extremely low prior achievement receives tutoring and no comparable student remains untreated.
Then the data provide little direct evidence about what would happen to such students without tutoring.
Positivity requires that relevant units have a nonzero probability of receiving the exposure alternatives being compared, conditional on the variables needed for the causal analysis.
In practice, researchers should examine whether exposure groups overlap sufficiently on important characteristics. Severe lack of overlap can make estimates depend heavily on modeling and extrapolation.
The exposure must be sufficiently well defined
Consider the question:
What is the effect of “using AI” on learning?
What intervention does that describe?
Using AI once to check grammar is not the same exposure as using an AI tutor for two hours each week. Different tools, purposes, durations, prompts, instructional contexts, and levels of human oversight could plausibly produce different outcomes.
A causal contrast becomes difficult to interpret when the exposure corresponds to many substantively different interventions.
Researchers should therefore define the treatment or exposure strategies precisely enough that the counterfactual comparison has a coherent meaning.
Measurement quality still matters after the causal structure is specified
A beautifully drawn causal diagram cannot rescue poorly measured variables.
If exposure is misclassified, important confounders are measured crudely, or outcome measurement differs systematically between exposure groups, the resulting causal estimate may be biased.
Observational causal inference therefore requires ordinary measurement validity alongside causal identification assumptions. Causal inference is not an exemption from measurement theory, despite occasional statistical enthusiasm to the contrary.
Selection can distort the causal comparison
Researchers may inadvertently condition on who enters the study, remains under follow-up, has complete data, or is included in the analysis.
If these selection processes depend on exposure, outcome determinants, or related variables, restricting analysis to the selected sample can produce bias.
Attrition is one example. If participants exposed to an intervention leave the study for reasons related to their potential outcomes, the remaining exposure groups may no longer support the intended comparison without additional assumptions and methods.
Statistical methods implement assumptions; they do not replace them
Regression adjustment, matching, propensity scores, inverse-probability weighting, standardization, g-methods, instrumental-variable approaches, regression discontinuity, difference-in-differences, and other methods can contribute to causal inference in suitable settings.
They do not provide interchangeable recipes for converting observational data into causal effects.
Each approach targets particular designs and relies on assumptions. The question should therefore be “Which design and assumptions make this method appropriate?” rather than “Which causal method can I run on this dataset?”
| Requirement |
Question to ask |
Why it matters |
| Defined causal contrast |
What exposure strategies are being compared? |
A causal effect has meaning only relative to an alternative. |
| Target population and estimand |
Whose causal effect and which effect are being estimated? |
Different estimands can answer different scientific or policy questions. |
| Temporal order |
Does exposure precede the relevant outcome? |
The proposed cause must occur before its effect. |
| Exchangeability |
Are exposure groups sufficiently comparable after justified adjustment? |
Confounding can make observed outcome differences differ from causal effects. |
| Positivity |
Are relevant exposure alternatives represented across the covariate patterns of interest? |
Without overlap, counterfactual comparisons may rely on unsupported extrapolation. |
| Consistency |
Does observed exposure correspond coherently to the intervention being defined? |
Ambiguous treatment versions can make the causal contrast unclear. |
| Valid measurement |
Are exposure, confounders, and outcomes measured adequately? |
Measurement error can distort causal estimates. |
| Selection |
Could inclusion, follow-up, or missingness depend on variables relevant to the causal effect? |
Selected samples can produce biased comparisons. |
Sensitivity analysis can test how conclusions depend on assumptions
Some assumptions required for observational causal inference cannot be verified directly from the data.
Sensitivity analyses can examine how conclusions would change under specified departures from assumptions, such as particular amounts of unmeasured confounding, alternative model specifications, exposure definitions, or missing-data mechanisms.
They do not prove that the primary assumptions are correct. Their value lies in showing whether the substantive conclusion is fragile or comparatively robust to plausible alternatives.
Triangulation can strengthen a causal argument
Evidence from different designs may have different sources of bias. If studies using distinct populations, measurement approaches, natural experiments, quasi-experimental strategies, or other designs converge on a similar conclusion despite different weaknesses, the combined evidence may support a stronger causal argument.
This is not a vote-counting exercise. Ten studies with the same confounding problem do not become causal merely through repetition.
The useful form of triangulation comes from evidence whose biases and assumptions differ in informative ways.
STROBE helps with reporting, not causal identification
Observational researchers are often advised to use STROBE, the Strengthening the Reporting of Observational Studies in Epidemiology statement.
That is appropriate for transparent reporting of cohort, case-control, and cross-sectional studies. But STROBE explicitly describes itself as guidance for reporting observational research rather than a prescription for study design or an instrument for evaluating study quality.
Following a reporting checklist does not establish that the causal assumptions are satisfied. Transparent reporting makes the causal argument easier to evaluate; it does not create the argument.
Sometimes the correct conclusion is still association
Not every observational dataset can answer the causal question you would like to ask.
If exposure timing is unclear, major confounders were not measured, the comparator represents a fundamentally different population, overlap is poor, or the exposure is too vaguely defined, a causal interpretation may require assumptions that are difficult to defend.
In those circumstances, describing the finding as an association is not methodological timidity. It is accurate calibration of the claim to the evidence.