Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Can an Observational Study Support a Causal Claim?

Observational data do not automatically restrict researchers to non-causal questions. Causal inference may be possible when the causal contrast, design, assumptions, timing, confounding strategy, and analysis are explicitly aligned.

44
Observational Studies and Causal Claims Guide 44 of 217
01 · The Question

If You Did Not Randomize the Exposure, Can You Still Study Its Causal Effect?

You want to know whether frequent use of an educational technology improves student achievement. Students already decide for themselves whether and how much to use it, so randomly assigning usage is not part of your study.

Does that mean you are limited to saying that technology use is “associated with” achievement, no matter how carefully the study is designed?

Not necessarily.

Observational studies can be designed and analyzed to address causal questions. Modern causal-inference frameworks explicitly provide methods for doing so. But the absence of random assignment changes what must be justified. The researcher must explain why the observed data, together with the design and assumptions, provide a credible comparison between the outcomes that would occur under alternative exposure conditions.

The difference between an observational association and an observational causal estimate is therefore not created by replacing the word “association” with “effect.” It comes from the causal question, design, identification assumptions, and analysis.

02 · The Short Answer

Observational Studies Can Support Causal Inference, but the Design Must Do Much More Work

In Brief

An observational study can support a causal claim when it is explicitly designed to estimate a well-defined causal effect and the assumptions needed to identify that effect from the observed data are sufficiently plausible.

This usually requires a clearly defined exposure or intervention contrast, appropriate temporal order, a credible strategy for addressing confounding and selection, adequate comparison or overlap between exposure conditions, valid measurement, and an analysis aligned with the causal estimand. Observational data alone do not make these assumptions true.

03 · What You Need to Know

“Observational” Describes How Exposure Occurred, Not Necessarily the Question You Are Allowed to Ask

In an observational study, researchers observe exposures or treatments that were not assigned through the study's randomization procedure. That distinction matters because the exposed and unexposed groups may differ systematically before the outcome occurs.

Yet observational research can still be organized around a causal estimand. Contemporary causal-inference frameworks explicitly use observational data to estimate causal effects when randomized trials are unavailable, infeasible, unethical, or otherwise unsuitable.

The difficult part is establishing why the observed comparisons identify the causal effect of interest.

Begin with a causal question, not with an available regression coefficient

A causal analysis should start by defining what effect you actually want to estimate.

Compare:

  • Are students who frequently use an AI tutoring system more likely to pass mathematics?
  • What would happen to mathematics pass rates if eligible students used the AI tutoring system regularly compared with if they did not use it?

The first asks about an observed association. The second defines a causal contrast between alternative exposure conditions.

The distinction matters because the second question requires researchers to reason about counterfactual outcomes. You need evidence about what outcomes would have occurred under exposure conditions different from those actually observed.

Specify the causal estimand before deciding how to estimate it

A causal effect is not simply “the effect of X on Y.” Researchers need to define the population, treatment or exposure strategies, outcome, relevant time horizon, and causal contrast.

Recent methodological frameworks for observational causal inference emphasize this explicit specification. One such framework organizes the problem around the target estimand, target population, target trial, and target validity.

That order is useful because an analytical method cannot tell you what scientific effect you intended to estimate. The estimand must come first.

The target trial provides a useful design benchmark

One influential approach asks researchers to describe the hypothetical randomized trial that would answer the causal question if such a trial were feasible.

This hypothetical experiment is called the target trial.

Researchers specify features such as:

  • eligibility criteria;
  • treatment or exposure strategies;
  • assignment procedures;
  • the start of follow-up;
  • outcomes;
  • the causal contrast of interest;
  • the analysis plan.

The observational study is then designed to emulate these components as closely as the available data permit.

Target trial emulation does not turn observational data into randomized data. Its value is conceptual and methodological: it makes the causal question explicit and can help researchers avoid design errors that arise when eligibility, treatment assignment, baseline, and follow-up are defined inconsistently.

Temporal order must support the proposed causal direction

A proposed exposure must occur before the outcome it is claimed to affect.

This can be difficult when observational data were collected without the causal question in mind. Exposure and outcome may be recorded concurrently, or an apparent baseline may occur after exposure has already begun.

Researchers should therefore establish the relevant temporal order rather than assuming that a longitudinal dataset automatically provides it.

The start of follow-up is particularly important. If eligibility, treatment classification, and follow-up begin at inconsistent times, researchers can introduce avoidable biases even when the subsequent statistical model is sophisticated.

Confounding is one of the central challenges

Without random assignment, exposure may be related to characteristics that also affect the outcome.

Suppose students who voluntarily use an academic-support platform are more motivated, have stronger prior achievement, or receive more instructor encouragement. Those characteristics could influence both platform use and later achievement.

If they are not appropriately addressed, the observed exposure-outcome difference may not represent the causal effect of the platform.

This is confounding.

The challenge is not simply to “control for some variables.” Researchers need to identify a sufficient set of pre-exposure confounders using substantive and causal knowledge and then measure and incorporate them appropriately.

More covariate adjustment is not necessarily better causal adjustment

A familiar observational analysis includes every variable available in the dataset and interprets the adjusted exposure coefficient as causal.

That strategy is not generally justified.

Variables can occupy different positions in the causal structure. Some are confounders. Others may be mediators caused by the exposure. Some can be colliders influenced by variables in the causal system. Adjusting indiscriminately can change the causal effect being estimated or introduce bias.

Variable selection should therefore follow the causal question and assumed causal structure rather than a rule such as “include everything associated with the outcome” or “adjust for everything significant in the univariate analysis.”

No unmeasured confounding is a strong assumption

Even careful adjustment can address only the relevant variables that were measured adequately.

For many observational causal analyses, researchers rely on an exchangeability assumption that, conditional on an appropriate set of measured variables, the exposure groups are sufficiently comparable with respect to their potential outcomes.

This assumption is demanding because important confounders may be unavailable, poorly measured, or unknown.

Crucially, a good balance table after matching or weighting cannot prove that no unmeasured confounding remains. The table describes measured information. The causal assumption extends beyond what is directly observed.

Positivity asks whether meaningful comparisons actually exist in the data

Suppose every student with extremely low prior achievement receives tutoring and no comparable student remains untreated.

Then the data provide little direct evidence about what would happen to such students without tutoring.

Positivity requires that relevant units have a nonzero probability of receiving the exposure alternatives being compared, conditional on the variables needed for the causal analysis.

In practice, researchers should examine whether exposure groups overlap sufficiently on important characteristics. Severe lack of overlap can make estimates depend heavily on modeling and extrapolation.

The exposure must be sufficiently well defined

Consider the question:

What is the effect of “using AI” on learning?

What intervention does that describe?

Using AI once to check grammar is not the same exposure as using an AI tutor for two hours each week. Different tools, purposes, durations, prompts, instructional contexts, and levels of human oversight could plausibly produce different outcomes.

A causal contrast becomes difficult to interpret when the exposure corresponds to many substantively different interventions.

Researchers should therefore define the treatment or exposure strategies precisely enough that the counterfactual comparison has a coherent meaning.

Measurement quality still matters after the causal structure is specified

A beautifully drawn causal diagram cannot rescue poorly measured variables.

If exposure is misclassified, important confounders are measured crudely, or outcome measurement differs systematically between exposure groups, the resulting causal estimate may be biased.

Observational causal inference therefore requires ordinary measurement validity alongside causal identification assumptions. Causal inference is not an exemption from measurement theory, despite occasional statistical enthusiasm to the contrary.

Selection can distort the causal comparison

Researchers may inadvertently condition on who enters the study, remains under follow-up, has complete data, or is included in the analysis.

If these selection processes depend on exposure, outcome determinants, or related variables, restricting analysis to the selected sample can produce bias.

Attrition is one example. If participants exposed to an intervention leave the study for reasons related to their potential outcomes, the remaining exposure groups may no longer support the intended comparison without additional assumptions and methods.

Statistical methods implement assumptions; they do not replace them

Regression adjustment, matching, propensity scores, inverse-probability weighting, standardization, g-methods, instrumental-variable approaches, regression discontinuity, difference-in-differences, and other methods can contribute to causal inference in suitable settings.

They do not provide interchangeable recipes for converting observational data into causal effects.

Each approach targets particular designs and relies on assumptions. The question should therefore be “Which design and assumptions make this method appropriate?” rather than “Which causal method can I run on this dataset?”

Requirement Question to ask Why it matters
Defined causal contrast What exposure strategies are being compared? A causal effect has meaning only relative to an alternative.
Target population and estimand Whose causal effect and which effect are being estimated? Different estimands can answer different scientific or policy questions.
Temporal order Does exposure precede the relevant outcome? The proposed cause must occur before its effect.
Exchangeability Are exposure groups sufficiently comparable after justified adjustment? Confounding can make observed outcome differences differ from causal effects.
Positivity Are relevant exposure alternatives represented across the covariate patterns of interest? Without overlap, counterfactual comparisons may rely on unsupported extrapolation.
Consistency Does observed exposure correspond coherently to the intervention being defined? Ambiguous treatment versions can make the causal contrast unclear.
Valid measurement Are exposure, confounders, and outcomes measured adequately? Measurement error can distort causal estimates.
Selection Could inclusion, follow-up, or missingness depend on variables relevant to the causal effect? Selected samples can produce biased comparisons.

Sensitivity analysis can test how conclusions depend on assumptions

Some assumptions required for observational causal inference cannot be verified directly from the data.

Sensitivity analyses can examine how conclusions would change under specified departures from assumptions, such as particular amounts of unmeasured confounding, alternative model specifications, exposure definitions, or missing-data mechanisms.

They do not prove that the primary assumptions are correct. Their value lies in showing whether the substantive conclusion is fragile or comparatively robust to plausible alternatives.

Triangulation can strengthen a causal argument

Evidence from different designs may have different sources of bias. If studies using distinct populations, measurement approaches, natural experiments, quasi-experimental strategies, or other designs converge on a similar conclusion despite different weaknesses, the combined evidence may support a stronger causal argument.

This is not a vote-counting exercise. Ten studies with the same confounding problem do not become causal merely through repetition.

The useful form of triangulation comes from evidence whose biases and assumptions differ in informative ways.

STROBE helps with reporting, not causal identification

Observational researchers are often advised to use STROBE, the Strengthening the Reporting of Observational Studies in Epidemiology statement.

That is appropriate for transparent reporting of cohort, case-control, and cross-sectional studies. But STROBE explicitly describes itself as guidance for reporting observational research rather than a prescription for study design or an instrument for evaluating study quality.

Following a reporting checklist does not establish that the causal assumptions are satisfied. Transparent reporting makes the causal argument easier to evaluate; it does not create the argument.

Sometimes the correct conclusion is still association

Not every observational dataset can answer the causal question you would like to ask.

If exposure timing is unclear, major confounders were not measured, the comparator represents a fundamentally different population, overlap is poor, or the exposure is too vaguely defined, a causal interpretation may require assumptions that are difficult to defend.

In those circumstances, describing the finding as an association is not methodological timidity. It is accurate calibration of the claim to the evidence.

04 · A Practical Example

From an Observed Difference to a Causal Question

Hypothetical Example

Does voluntary AI tutoring improve course completion?

A university has records for 8,000 students. Some voluntarily used an AI tutoring system during the semester, while others did not. Students who used it had a higher course-completion rate.

Association analysis The researchers compare users with non-users and find that users are more likely to complete the course. This establishes an observed association, not yet the causal effect of tutoring.
Causal question The researchers define the target effect as the difference in course completion if eligible students regularly used the tutoring system compared with if those same eligible students did not use it during the semester.
Design problem Tutoring use was self-selected. Users differ from non-users in prior achievement, course engagement, attendance, program of study, and several other measured characteristics. Motivation, however, is not measured well.
Causal analysis The researchers emulate the relevant aspects of a target trial, establish a common start of follow-up, adjust for justified pre-exposure confounders using an appropriate method, examine overlap, assess alternative specifications, and conduct sensitivity analyses for assumptions they cannot verify directly.

The analysis is now organized around a causal question rather than merely an adjusted association. Whether the resulting causal estimate is credible depends on how plausible the assumptions are, particularly whether important residual confounding remains.

The phrase “observational study” neither proves nor disproves the causal interpretation. The causal argument has to be earned through the design.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Causal Inference From Observational Data

Misconception

Observational Studies Can Never Support Causal Claims

That is too restrictive. Modern causal-inference methods explicitly address causal questions using observational data. The causal interpretation depends on the design, estimand, assumptions, measurement, analysis, and credibility of the counterfactual comparison rather than on the observational label alone.

Misconception

Adjusting for Several Covariates Makes an Observational Result Causal

No. Adjustment is useful only when the selected variables and method correspond to the causal structure and relevant assumptions. Important confounders may remain unmeasured, while inappropriate adjustment can introduce bias or change the effect being estimated.

Misconception

Propensity Score Matching Recreates Randomization

No. Matching can improve comparability on observed information under appropriate implementation. It does not reproduce the treatment-assignment mechanism of a randomized experiment and cannot guarantee balance on unmeasured confounders.

Misconception

A Longitudinal Observational Study Is Automatically Causal

No. Longitudinal data can establish important temporal information, but confounding, selection, measurement error, attrition, and other sources of bias can remain.

Misconception

Target Trial Emulation Turns Observational Data Into a Randomized Trial

No. It uses the protocol of a hypothetical randomized trial as a framework for defining and organizing an observational causal analysis. Treatment remains observational, so causal identification still depends on assumptions that random assignment would otherwise help establish by design.

Misconception

Calling a Coefficient an “Effect” Makes It a Causal Estimate

No. Causal terminology should correspond to a causal estimand identified under a defensible design and assumptions. A regression coefficient from an associational model does not acquire causal meaning through wording.

06 · What This Means for You

Decide Whether You Have a Causal Design Before Writing a Causal Conclusion

If you are working with observational data, do not begin by asking whether your software offers propensity scores, inverse-probability weighting, or another method with “causal” in the documentation.

Begin with the causal question.

A simple decision framework

If you only want to describe whether exposure and outcome are related
An associational analysis may be entirely appropriate. Do not impose causal machinery that the research question does not require.
If you want to estimate what would happen under alternative exposure conditions
Define the causal estimand, target population, exposure strategies, outcome, follow-up, and relevant counterfactual contrast explicitly.
If important confounders are known and measured adequately
Choose a design and analytical strategy that addresses those confounders under clearly stated assumptions and assess whether sufficient overlap exists.
If a major confounder is unavailable or measured poorly
Determine whether another design, data source, identification strategy, or sensitivity analysis can address the problem; otherwise constrain the causal claim.
If the causal assumptions are difficult to defend
Report the observed relationship as an association rather than presenting a stronger causal interpretation than the design supports.
Watch Out

Do not treat causal inference as a statistical upgrade applied after an observational study has already been designed. Eligibility, baseline, treatment definition, temporal order, confounder measurement, follow-up, and outcome assessment often need to be planned around the causal question before the final model is fitted.

If you are unsure whether your current design has crossed the line from describing a relationship to estimating a causal effect, examine whether the design supports association but not causation before choosing causal language.

07 · A Quick Checklist

Before Making a Causal Claim From Observational Data

Before interpreting an observational estimate causally, check:
Have you defined the causal effect you want to estimate rather than merely the variables you want to associate?
Are the exposure or intervention alternatives sufficiently well defined?
Does the exposure precede the relevant outcome, with a clearly defined start of follow-up?
Have you identified plausible pre-exposure confounders using substantive and causal knowledge?
Were the important confounders measured with sufficient quality for the intended analysis?
Is there adequate overlap between exposure groups for the population and contrast you intend to study?
Could selection, attrition, missing data, or differential measurement distort the comparison?
Can you state the identifying assumptions required for your causal interpretation?
Have you examined how sensitive the conclusion is to plausible violations of assumptions that cannot be verified directly?
Does the language of the conclusion match the causal effect the design actually identifies?
08 · Frequently Asked Questions

Frequently Asked Questions About Causal Claims From Observational Studies

Can observational studies prove causation?

Scientific evidence is rarely usefully described in terms of absolute proof. Observational studies can support causal inference when their design and assumptions identify a well-defined causal effect, but uncertainty about confounding, selection, measurement, and other assumptions should be represented honestly.

Do I need a randomized controlled trial to make any causal claim?

No. Randomized trials provide a particularly strong treatment-assignment mechanism, but many causal questions cannot be randomized. Observational causal inference uses other designs and assumptions to identify causal effects, although those assumptions may be more demanding.

Does regression adjustment make an observational study causal?

No. Regression can implement adjustment for measured confounders, but causal interpretation depends on whether the relevant confounders were correctly identified and measured, the model and estimand are appropriate, sufficient overlap exists, and other identifying assumptions are plausible.

Does propensity score matching eliminate confounding?

It can improve balance on measured covariates under suitable implementation. It does not guarantee control of unmeasured confounding and should not be treated as equivalent to random assignment.

What is target trial emulation?

It is an approach that specifies the hypothetical randomized trial that would answer the causal question and then attempts to emulate its key protocol components using observational data. It can clarify the estimand and reduce avoidable design biases, but it does not make observational treatment assignment random.

Can a cross-sectional observational study support causation?

Concurrent measurement often makes causal interpretation difficult because temporal ordering may be unclear, but the answer depends on the specific variables, timing information, design, and identification strategy. The label “cross-sectional” alone is not a complete causal assessment.

What if I cannot rule out unmeasured confounding?

That is common in observational causal research. Researchers should make the assumption explicit, use design strategies and sensitivity analyses where appropriate, consider complementary evidence, and calibrate the strength of the causal conclusion to how plausible the assumption is.

Does following STROBE mean my observational study can make causal claims?

No. STROBE is a reporting guideline for observational studies and explicitly states that its recommendations are not prescriptions for study design or an instrument for evaluating study quality. Transparent reporting helps readers evaluate a causal analysis, but it does not establish causal identification.

09 · The Bottom Line

Observational Does Not Mean Non-Causal, but Causation Must Be Designed Into the Study

The Bottom Line

An observational study can support a causal claim when it targets a clearly defined causal effect and the design, measurement, comparison, analysis, and identifying assumptions provide a credible basis for estimating that effect.

Do not assume that observational data prohibit causal inference, but do not assume that adjustment creates it either. Define the counterfactual contrast first, establish the relevant timing, address confounding and selection explicitly, examine overlap and measurement, and make the strength of the conclusion proportional to the assumptions the study requires.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes