Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should You Look for When a Paper Makes a Causal Claim?

A causal claim says more than two variables are associated. It says changing one would change the other. Evaluate whether the study establishes temporal order, provides a credible comparison, addresses confounding and selection, measures the relevant variables adequately, and rules out plausible alternative explanations.

153
How to Evaluate a Causal Claim Guide 153 of 247
01 · The Question

What Does a Study Need Before You Can Say That X Caused Y?

A paper finds that students who use generative AI more frequently earn higher grades.

Did AI use improve their grades?

Perhaps. But perhaps students who use AI frequently are already more academically motivated. Perhaps stronger students know how to use the technology more effectively. Perhaps particular instructors encourage both AI use and better study practices. Perhaps students with higher grades simply have more demanding academic work and therefore use AI more often.

An association tells you that two things vary together. A causal claim goes further. It asks what would happen to the outcome if the exposure, intervention, policy, behavior, or condition were changed while the relevant alternatives were otherwise comparable.

That is a much more demanding question.

When a paper makes a causal claim, the central task is to ask whether the study has created a credible comparison between what happened and what would plausibly have happened under the relevant alternative.

02 · The Short Answer

Look for a Credible Counterfactual, Not Merely a Significant Association

In Brief

When a paper makes a causal claim, check whether the proposed cause precedes the outcome, whether the comparison groups provide a credible estimate of what would have happened otherwise, whether important confounding and selection processes have been addressed, whether exposure and outcome were measured appropriately, and whether the analysis matches the causal question and study design.

Randomization can make causal inference more credible when implemented and analyzed appropriately, but even randomized trials require critical appraisal. Observational studies can also investigate causal effects, although their conclusions depend more heavily on assumptions about confounding, selection, measurement, positivity, model specification, and other sources of bias.

03 · What You Need to Know

Causal Inference Is About What Would Have Happened Otherwise

First, identify whether the paper is actually making a causal claim

Authors do not always write the word cause.

Causal claims can appear through verbs such as:

  • increases;
  • reduces;
  • improves;
  • prevents;
  • leads to;
  • results in;
  • produces;
  • affects;
  • influences;
  • drives.

A statement such as "frequent AI use improves academic performance" is causal because it implies that changing AI use would change performance.

Compare that with:

Frequent AI use was associated with higher academic performance.

The second statement describes a relationship without necessarily claiming that changing one variable would change the other.

This distinction is central when checking whether authors have made their conclusion stronger than their evidence supports.

A causal question is different from an associational question

Suppose researchers observe that students who attend optional tutoring sessions earn higher examination scores.

An associational question asks:

Do students who attend tutoring have different examination scores from students who do not?

A causal question asks something closer to:

For comparable students, what would happen to examination scores if they received tutoring rather than not receiving it?

The difference is subtle but fundamental.

Association What differences or relationships are observed among people who happened to have different exposures or experiences?
Causal effect How would the outcome differ if the relevant exposure or intervention were changed under otherwise comparable conditions?

A study can estimate an association accurately while still providing a biased estimate of a causal effect.

The counterfactual is the outcome you cannot observe directly

Causal reasoning involves comparison with a counterfactual.

Imagine one student receives an AI tutoring intervention and earns a score of 85.

To know the causal effect for that student, you would ideally also know what the same student's score would have been at the same time under the same circumstances without the AI tutor.

You cannot observe both outcomes simultaneously.

This is the fundamental problem of causal inference. Research designs attempt to construct a credible substitute for the unobserved counterfactual.

Randomized experiments do this by creating groups intended to be comparable before intervention assignment. Observational causal studies use design and analytical strategies intended to make exposed and unexposed groups sufficiently comparable under explicit assumptions.

Ask whether the cause occurred before the effect

Temporality is indispensable to causal interpretation.

If X causes Y, X must occur before the relevant change in Y.

This sounds obvious, yet many studies measure exposure and outcome at the same time.

Suppose a cross-sectional survey finds that students reporting heavier social-media use also report greater anxiety. Which came first?

Perhaps social-media use contributed to anxiety. Perhaps anxious students use social media more frequently. Perhaps both are influenced by another factor.

A one-time measurement often cannot establish temporal ordering adequately.

Watch Out

If exposure and outcome were measured at the same time, inspect causal language carefully. A cross-sectional association can be real while the direction of causation remains unresolved.

Longitudinal evidence helps with temporality but does not solve causality by itself

Following people over time can establish that an exposure was measured before a later outcome.

That is stronger evidence about temporal order.

It does not automatically eliminate confounding, selection bias, measurement problems, or other alternative explanations.

Suppose students who report frequent AI use in September earn higher grades in December. AI use preceded the measured grades, but frequent users may already have differed in prior achievement, motivation, course difficulty, socioeconomic circumstances, or instructor support.

Temporal precedence is necessary for many causal claims. It is not sufficient.

Randomization is powerful because it targets comparability

In a well-conducted randomized experiment, assignment to intervention conditions is determined by a random mechanism rather than participant choice, clinician judgment, teacher preference, or another process related to prognosis.

With adequate implementation, randomization can help balance both measured and unmeasured baseline causes of the outcome across groups, apart from chance variation.

This makes differences in outcomes more plausibly attributable to the intervention.

But the word "randomized" is not enough. You still need to ask:

  • Was the allocation sequence genuinely random?
  • Was allocation concealed before assignment?
  • Did participants remain in their assigned groups?
  • Were there important deviations from the intended intervention?
  • Was outcome measurement comparable across groups?
  • Was attrition substantial or differential?
  • Was analysis consistent with the randomized design?
  • Were outcomes selectively reported?

Randomization strengthens causal inference when its logic survives the conduct and analysis of the study.

Random assignment and random sampling solve different problems

A randomized trial does not need a nationally representative sample to estimate a causal effect within the study population.

Random assignment concerns comparability between intervention groups. Random sampling concerns how units are selected from a broader population.

Random assignment Primarily strengthens internal causal comparison between conditions.
Random sampling Can strengthen inference from the sample to a defined target population.

A trial can therefore provide a credible causal effect among the studied participants while leaving questions about whether the same effect applies elsewhere.

This is one reason sample appropriateness and causal validity should be evaluated separately.

Ask whether the comparison group represents a credible alternative

Causal inference depends heavily on the comparator.

Suppose students using an optional AI tutor outperform students who do not use it.

Those groups may differ before the tutoring begins.

Users might be more motivated, more digitally skilled, more likely to seek help, enrolled in different courses, or encouraged by different instructors.

If those characteristics also affect academic performance, the observed group difference does not isolate the effect of the AI tutor.

The central question is:

Would these groups have had similar outcomes if their exposure had been the same?

In observational research, that question cannot usually be answered directly. It depends on design choices, measured covariates, assumptions, and analytical methods.

Confounding is an alternative explanation built into the comparison

A confounder is not merely any variable correlated with both exposure and outcome.

In causal reasoning, confounding arises when the exposed and comparison groups differ in causes of the outcome in ways that distort the causal contrast of interest.

Consider coffee consumption and academic performance.

If students who study longer drink more coffee and studying longer improves grades, study time could contribute to an observed association between coffee and grades even if coffee itself has little causal effect.

Observed association Students drinking more coffee have higher grades.
Potential confounder Students who study longer may drink more coffee.
Independent influence Longer study time may itself improve grades.
Causal problem Part of the coffee-grade association may therefore reflect differences in study behavior rather than coffee.

The practical task is to identify plausible common causes of the exposure and outcome that could distort the comparison.

Do not identify confounders solely by statistical significance

Researchers sometimes select covariates because they are significantly associated with the outcome or because adding them changes a coefficient.

Modern causal inference places greater emphasis on substantive causal knowledge.

Variables should be considered according to their role in the causal structure rather than merely whether they satisfy a data-driven statistical criterion.

A variable can be an important confounder without showing a statistically significant association in one sample. Conversely, adjusting for some statistically associated variables can introduce bias.

When a causal claim is central, ask why the adjustment variables were selected.

More adjustment does not necessarily mean less bias

A regression model with twenty control variables can look reassuring.

It may be appropriate. It may also include variables that should not be adjusted for.

For example, if a variable lies on the causal pathway between exposure and outcome, adjusting for it may remove part of the effect the researchers intend to estimate.

Other variables can behave as colliders, where conditioning on them creates an association between otherwise unrelated causes and introduces bias.

The technical details can become complex, but the critical-reading principle is straightforward:

Adjustment variables should be chosen according to the causal question and assumed causal structure, not according to the principle that controlling for more variables is always safer.

A causal diagram can make the authors’ assumptions visible

Directed acyclic graphs, or DAGs, are increasingly used to represent assumptions about causal relationships among variables.

A DAG can help researchers identify potential confounders, mediators, colliders, and variables that need not or should not be adjusted for under a particular causal model.

The graph does not prove that the assumed causal structure is correct. Its value is partly transparency: it forces researchers to make assumptions visible enough to critique.

If a paper makes a sophisticated observational causal claim, a clearly articulated causal model can be more informative than a statement that the researchers "controlled for all relevant variables."

Ask whether important confounders were measured well enough

Including a variable in a regression model does not guarantee that confounding by that variable has been eliminated.

Suppose motivation is an important confounder, but the study measures it using one vague self-report item. Residual confounding may remain because the variable is measured imperfectly.

The same issue applies when broad categories inadequately capture complex characteristics such as socioeconomic status, disease severity, prior achievement, or baseline health.

Causal inference therefore depends not only on which variables were measured but how well those variables were measured.

Unmeasured confounding cannot be removed by ordinary regression

This point is simple but consequential.

If an important confounder was never measured, a conventional regression model cannot adjust for it by including other unrelated variables.

Authors may argue that unmeasured confounding is unlikely, use negative controls, instrumental variables, quantitative bias analysis, sensitivity analyses, natural experiments, or other strategies depending on the problem.

Those approaches can strengthen causal reasoning under their own assumptions.

But "adjusted analysis" should never be read as synonymous with "all confounding eliminated."

Selection can create associations that are not causal

Selection bias occurs when inclusion in the study or analysis depends on variables in a way that distorts the relationship of interest.

This can occur during recruitment, follow-up, exclusion, complete-case analysis, or conditioning on particular subsets of participants.

Suppose a study examines whether a demanding intervention improves performance but analyzes only participants who completed the full intervention. Completion may itself depend on motivation, baseline ability, tolerance of the intervention, or early response.

The resulting completers may no longer provide a credible comparison for the causal effect among everyone who started.

Ask how participants entered the analytical sample and whether selection could depend jointly on exposure and outcome-related factors.

Attrition can break a randomized comparison

Randomization creates comparable groups at assignment. Differential loss after assignment can erode that advantage.

Imagine 10% of control participants but 45% of intervention participants are missing at follow-up because participants who found the intervention ineffective were more likely to withdraw.

The remaining groups may no longer represent the randomized comparison cleanly.

Cochrane's Risk of Bias 2 framework explicitly evaluates missing outcome data as one domain when assessing randomized trials.

Look beyond the initial randomization and examine who actually contributed outcome data.

Ask whether the exposure or intervention was defined clearly enough

Causal claims require a reasonably clear idea of what is being changed.

"AI use" may include brainstorming, editing, automated feedback, coding, summarization, translation, answer generation, or many other activities.

If one student uses AI once per month for grammar correction and another uses it daily to generate assignment drafts, treating both simply as "AI users" creates an ambiguous exposure.

A causal question should make the intervention or exposure sufficiently well defined that you can understand what hypothetical change the causal claim refers to.

If authors conclude that "AI use improves learning," ask: what form of AI use, how much, for whom, under what conditions, compared with what?

Ask whether the comparison condition is equally clear

Causal effects are contrasts.

An intervention does not have one universal effect independent of what it is compared with.

An AI tutor compared with no tutoring may produce one effect. Compared with an expert human tutor, it may produce another. Compared with conventional digital exercises, another again.

Comparison Causal question
AI tutor vs. no additional support What is the effect of providing the AI tutor rather than nothing additional?
AI tutor vs. human tutoring What is the effect of substituting AI tutoring for human tutoring?
AI tutor vs. standard digital practice What is the effect relative to the existing digital alternative?

A paper saying an intervention "works" should prompt the question: compared with what?

Baseline differences deserve attention, but not every difference proves randomization failed

Randomized groups can differ at baseline by chance, particularly in small trials.

The existence of some imbalance does not automatically demonstrate defective randomization.

More important questions include whether randomization and allocation concealment were implemented properly, whether major prognostic imbalance exists, and whether the planned analysis appropriately incorporates relevant baseline measurements.

In nonrandomized studies, baseline differences have a different meaning because treatment or exposure selection itself may be related to prognosis.

Ask whether treatment actually differed between the groups

Assignment to an intervention does not guarantee that participants received or adhered to it.

Some participants may cross over, discontinue treatment, receive additional interventions, or fail to engage with the assigned condition.

This affects the causal question.

An intention-to-treat analysis in a randomized trial estimates the effect of assignment to an intervention strategy under the trial's conditions. A per-protocol effect asks a different question about adherence to the intervention and requires additional assumptions because adherence is not randomized.

Do not treat different causal estimands as interchangeable merely because they appear in the same trial.

Ask what causal effect the study is actually estimating

The phrase "the treatment effect" can conceal several different questions.

Researchers might estimate:

  • the effect of being assigned to an intervention;
  • the effect of receiving the intervention;
  • the effect among people who adhered;
  • the average effect in the whole study population;
  • the effect among people who actually received treatment;
  • the effect under a particular intervention strategy over time.

These quantities are related but not identical.

Modern causal inference often emphasizes specifying the causal estimand clearly before choosing the analytical method.

If you cannot tell what comparison the estimate represents, causal interpretation becomes difficult.

Observational causal studies should make the hypothetical intervention clear

One useful strategy is to imagine the randomized trial that the observational study is trying to emulate.

Hernán and colleagues have developed the target trial framework precisely for this purpose. Researchers specify features such as eligibility criteria, treatment strategies, assignment procedures, follow-up, outcomes, causal contrast, and analysis, then attempt to emulate that trial using observational data.

This does not make observational data equivalent to randomized data. It can clarify the causal question and expose design choices that otherwise remain implicit.

When reading an observational causal study, ask:

What hypothetical intervention and comparison is this study trying to estimate?

If that question has no clear answer, the causal claim may be poorly defined.

Time zero should align across eligibility, treatment, and follow-up

Causal analyses can be biased when the start of eligibility, treatment assignment, and outcome follow-up do not align appropriately.

For example, suppose patients are classified as "treated" only if they survive long enough to receive treatment, while follow-up begins earlier. The period during which they must survive to become classified as treated can create immortal time bias.

The details vary across fields, but the general question is accessible:

When does the causal comparison begin, and are exposure classification and outcome follow-up aligned from that point?

Temporal alignment is especially important in longitudinal observational research.

Reverse causation should be considered explicitly

Sometimes the proposed outcome can influence the proposed exposure.

Suppose poor academic performance is associated with greater use of AI tools. Perhaps AI use affects performance. Perhaps struggling students turn to AI more frequently because their performance is already poor.

Cross-sectional designs are particularly vulnerable to this ambiguity, but reverse causation can also complicate longitudinal studies when exposure measurement occurs after the underlying outcome process has already begun.

Ask whether the study design separates the proposed cause from the developing outcome sufficiently to make the claimed direction credible.

Measurement can create an apparent causal effect

Suppose participants know whether they received an intervention and the outcome is subjective.

If intervention participants expect improvement, their self-reports may change partly because of expectations rather than the underlying construct the researchers intend to measure.

Likewise, assessors who know treatment status may rate outcomes differently.

Blinding can reduce some forms of differential measurement, although it is not feasible in every study.

The causal question is not merely whether the groups differ, but whether the difference reflects the intended causal pathway rather than the measurement process itself.

Mediation is not established merely because the mediator changed

A study finds that an intervention improves achievement and also increases motivation. The authors conclude that motivation explains the achievement effect.

That conclusion requires more than observing both changes.

To establish mediation, researchers need a defensible causal model relating intervention, mediator, outcome, and relevant confounders, along with temporal and analytical assumptions appropriate to the mediation estimand.

Post-treatment confounding can make mediation analysis particularly challenging.

A plausible mechanism is not the same thing as an established causal pathway.

Dose-response patterns can strengthen a causal argument but do not prove causation

If greater exposure is associated with progressively larger changes in an outcome, the pattern may be compatible with causality.

Bradford Hill included biological gradient among his well-known viewpoints for considering causal explanations. He explicitly did not present his viewpoints as hard-and-fast criteria capable of proving causation. Later causal-inference scholarship likewise cautions against using them as a mechanical checklist.

A dose-response relationship can still arise through confounding, selection, measurement, or other processes.

Use it as one piece of a causal argument, not as a causal certificate.

Consistency across studies strengthens evidence but is not sufficient by itself

If several independent studies using different populations and methods produce similar findings, causal interpretation may become more credible.

But repeated bias can also produce repeated findings.

Ten observational studies relying on similar self-selected samples and the same poorly measured confounders do not necessarily provide ten independent solutions to the same methodological problem.

Look for triangulation across designs, data sources, populations, measurements, and analytical approaches whose biases differ.

Plausibility is useful but can become hindsight

A causal explanation that fits established theory or biological knowledge may be more credible than one that contradicts well-supported knowledge without explanation.

But plausibility depends on what is currently known.

Researchers can often construct a plausible story after seeing almost any interesting association.

Bradford Hill treated plausibility as one consideration rather than a necessary criterion, and contemporary causal assessment generally retains that caution.

Mechanistic plausibility should complement strong design and evidence, not replace them.

Do not use the Bradford Hill viewpoints as a scorecard

Strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy are historically influential considerations in causal assessment.

Hill himself did not describe them as rules that had to be satisfied before causality could be accepted. Modern reviews similarly emphasize that the viewpoints should not be used as a mechanical checklist in which accumulating enough ticks proves causation.

Some are more informative in particular contexts than others. Temporality, for example, has a logical role that analogy does not.

Use such considerations to organize thinking about a body of evidence, while still evaluating the causal design and assumptions of the specific study.

A strong association is not necessarily causal

Large associations can be harder to explain entirely through modest confounding, but strength alone does not prove causation.

Selection bias, severe confounding, measurement problems, or other mechanisms can sometimes produce large associations.

Conversely, a genuinely causal effect can be small.

Do not equate effect magnitude with causal validity.

Statistical significance tells you almost nothing about causality by itself

A causal estimate can have p <.001 and still be biased.

Statistical significance addresses sampling variation under a specified model. It does not establish that:

  • the groups were comparable;
  • confounding was controlled adequately;
  • selection bias was absent;
  • the causal direction was correct;
  • the exposure and outcome were measured validly;
  • the model was correctly specified;
  • the causal assumptions were satisfied.

Likewise, a non-significant estimate does not prove absence of a causal effect.

When reading the results, focus on the estimated effect, its uncertainty, and what the underlying design permits you to infer.

Sensitivity analyses can reveal how fragile a causal conclusion is

Observational causal analyses often depend on assumptions that cannot be fully verified from the observed data.

Sensitivity analyses can examine how results change under alternative specifications, definitions, missing-data assumptions, adjustment sets, or assumptions about unmeasured confounding.

If modest plausible changes overturn the conclusion, the causal finding is fragile.

If results remain similar across several defensible approaches, confidence may increase, although robustness to analyzed alternatives does not prove that all relevant biases have been eliminated.

Negative controls can test some alternative explanations

Some studies use negative-control exposures or outcomes that should not be causally affected under the proposed mechanism but may share similar sources of bias.

An unexpected association with a negative control can suggest residual confounding, selection, or measurement problems.

Negative controls are not universal requirements and depend on finding appropriate controls. When used well, they can provide additional information about whether an apparent causal effect might reflect bias.

Instrumental variables and natural experiments rely on strong assumptions too

When randomization is unavailable, researchers may exploit natural experiments or instrumental variables to approximate causal contrasts.

These methods can be powerful when their assumptions are credible.

An instrumental variable, for example, must affect the outcome through the exposure pathway relevant to the causal question and satisfy other conditions that are often difficult to verify completely.

Do not treat labels such as "natural experiment," "instrumental variable," "regression discontinuity," or "difference-in-differences" as automatic guarantees of causality.

Ask what feature of the design creates the quasi-experimental comparison and what assumptions are necessary for that comparison to identify the causal effect.

Difference-in-differences depends on more than having before-and-after data

Difference-in-differences designs compare changes over time between exposed and comparison groups.

A central assumption is that, absent the intervention, the groups would have followed sufficiently comparable outcome trends for the intended causal contrast.

Simply observing two groups before and after a policy does not guarantee that assumption.

Look for evidence about pre-intervention trends, concurrent events, composition changes, treatment timing, and sensitivity analyses where appropriate.

Regression discontinuity depends on what happens around the cutoff

Regression discontinuity designs can provide compelling causal evidence when treatment assignment changes at a threshold and units near that threshold are otherwise sufficiently comparable.

Critical questions include whether the assignment rule was followed, whether participants could manipulate the running variable around the cutoff, whether other interventions changed at the same threshold, and whether the modeled relationship around the cutoff is appropriate.

The resulting causal effect is often most directly interpretable near the threshold rather than automatically across the entire population.

Interrupted time-series designs need a credible counterfactual trend

An outcome changing after a policy or intervention does not automatically mean the policy caused the change.

The outcome may already have been changing, another event may have occurred simultaneously, seasonality may matter, or measurement practices may have changed.

Interrupted time-series analyses attempt to distinguish intervention-related changes from the underlying trend using repeated observations before and after the interruption.

The credibility of the causal interpretation depends on the number and quality of observations, model specification, concurrent events, and whether the pre-intervention trend provides a reasonable basis for the counterfactual.

Observational causal inference is possible, but its assumptions should be visible

It is too simplistic to say that observational studies can never support causal inference.

Modern causal inference provides frameworks and methods for estimating causal effects from observational data when the causal question is clearly specified and assumptions are sufficiently credible.

Recent methodological guidance emphasizes, however, that observational causal estimates depend on assumptions such as exchangeability, positivity, and appropriate measurement and modeling that can be difficult to verify completely.

The appropriate response is neither "observational means no causality" nor "advanced causal model means causality established."

Ask what assumptions connect the observed data to the causal effect and whether those assumptions are plausible.

Reporting guidelines help you see what you need, but they do not certify causality

For observational research, STROBE recommends transparent reporting of study design, participant selection, variables, measurement, efforts to address bias, statistical methods, confounding adjustment, missing data, subgroup analyses, sensitivity analyses, and limitations.

That information makes critical appraisal possible.

STROBE itself explicitly notes that its checklist is intended to improve reporting and is not an instrument for evaluating study quality.

A well-reported observational study can still have weak causal identification. Poor reporting can make an otherwise thoughtful causal design difficult to evaluate.

External validity comes after identifying a credible effect

Suppose a randomized trial provides strong evidence that an intervention caused improvement among the studied participants.

A second question remains:

Would the same causal effect occur in other populations, settings, implementations, or time periods?

Internal causal validity and generalizability are distinct.

A tightly controlled trial can estimate an internally credible effect while leaving uncertainty about routine implementation elsewhere.

Do not broaden a credible causal effect beyond the populations and contexts the evidence can reasonably inform.

One study rarely settles a major causal question

Causal confidence often develops from a body of evidence rather than one paper.

Different designs can contribute complementary information. Randomized trials may provide strong intervention comparisons but involve restricted populations. Observational studies may examine longer-term outcomes or real-world settings. Mechanistic research can clarify pathways. Natural experiments may provide evidence where randomized intervention is impossible.

When different approaches with different biases converge on a similar conclusion, the causal case can become stronger.

This is particularly important for exposures that cannot ethically or practically be randomized.

04 · A Practical Example

Does Generative AI Improve Academic Performance?

Hypothetical Example

A causal claim from observational data

Imagine a longitudinal study of 4,000 university students. At the beginning of the semester, students report how frequently they use generative AI for academic work. At the end of the semester, researchers obtain their course grades. Students reporting frequent AI use earn higher grades, and the association remains after adjustment for age, sex, prior GPA, and study time. The paper concludes that generative AI use improves academic performance.

1. Identify the causal claim "Improves" means the paper is claiming that changing AI use would change academic performance.
2. Check temporality AI use was measured before the end-of-semester grades. This is stronger than measuring both simultaneously, although students may already have been using AI and experiencing academic difficulties or advantages before baseline.
3. Examine treatment assignment Students chose whether and how frequently to use AI. Exposure was not randomized.
4. Identify plausible confounding Prior GPA and study time were measured, but digital competence, motivation, course difficulty, instructor practices, socioeconomic resources, and reasons for using AI may still differ between frequent and infrequent users.
5. Examine measurement AI use is self-reported as frequency. The measure may not distinguish productive uses such as feedback from uses that could affect learning differently.
6. Examine adjustment Regression adjustment addresses measured covariates included in the model. It does not demonstrate that unmeasured confounding is absent.
7. Consider alternative explanations Higher-performing or more digitally skilled students may use AI differently. Instructor practices could influence both AI use and grades.
8. Form a proportionate conclusion Frequent reported AI use preceded and was associated with higher end-of-semester grades after adjustment for several measured characteristics. The observational design strengthens temporal interpretation but does not by itself establish that increasing AI use would cause grades to improve.

The study may contribute useful evidence to a causal question. The issue is whether the evidence has crossed the much higher threshold required for the paper's definitive causal wording.

05 · What Researchers Often Get Wrong

Common Mistakes When Evaluating Causal Claims

Misconception

Correlation never tells us anything about causation

Associations are often part of causal investigation. The problem is treating association alone as sufficient evidence of causality. Causal inference requires additional design features, assumptions, temporal information, comparison logic, and evidence addressing plausible alternatives.

Misconception

Only randomized trials can ever support causal claims

Randomization provides a particularly powerful basis for many causal comparisons, but observational designs can also investigate causal effects using explicit causal frameworks and defensible assumptions. Their causal interpretation generally requires greater attention to confounding, selection, measurement, and identification assumptions.

Misconception

A longitudinal study proves causation because X came before Y

No. Temporal precedence addresses one essential requirement, but confounding, selection, measurement, and other alternative explanations can remain. Longitudinal evidence is stronger than simultaneous measurement for temporal questions but is not automatically causal.

Misconception

If researchers controlled for confounders, the estimate is causal

Not automatically. The result depends on whether the relevant confounders were identified, measured adequately, and adjusted appropriately. Unmeasured and residual confounding may remain, and adjusting for inappropriate variables can introduce bias.

Misconception

A very small p-value makes causation more likely

A small p-value can indicate that the observed estimate would be unusual under a specified null model. It does not tell you whether the comparison is confounded, selected, mismeasured, temporally ambiguous, or otherwise biased. Statistical significance and causal validity are different questions.

Misconception

A dose-response relationship proves causation

No. A gradient can strengthen a causal argument in some settings, but confounding or other biases can also produce dose-response patterns. Bradford Hill's viewpoints were intended as considerations, not a checklist whose completion proves causality.

Misconception

If several studies find the same association, the relationship must be causal

Consistency can increase confidence, especially when studies use different designs and sources of bias. But repeated studies can share the same methodological weaknesses. Consider whether evidence converges across genuinely different approaches rather than simply counting similar findings.

06 · What This Means for You

Reconstruct the Causal Comparison Before Accepting the Claim

You do not need to solve every philosophical problem of causation whenever a paper uses the word "affects." Focus on the causal comparison the study needs to support.

A simple decision framework

First: What exactly is the proposed cause?
Define the exposure, intervention, treatment, behavior, or policy precisely enough to know what is supposedly being changed.
Next: What is the outcome?
Check that it was measured adequately and at a time compatible with the causal direction being claimed.
Next: What is the counterfactual comparison?
Ask what would have happened to comparable units under the relevant alternative exposure or intervention.
Next: How were the comparison groups created?
Determine whether exposure was randomized, naturally occurring, self-selected, policy-determined, threshold-based, or produced through another mechanism.
Next: What could make the groups differ besides the proposed cause?
Identify plausible confounders, selection processes, concurrent interventions, baseline differences, and reverse causation.
Next: How were those alternatives addressed?
Examine design restrictions, randomization, matching, adjustment, weighting, quasi-experimental features, sensitivity analyses, negative controls, or other methods relevant to the causal question.
Next: What assumptions remain?
Identify assumptions that cannot be verified completely from the observed data and consider how plausible they are in the substantive context.
Finally: How strong should the causal language be?
Use language proportionate to how convincingly the design and evidence rule out important noncausal explanations.

A useful final question is: What plausible process other than X causing Y could have produced the observed difference, and what did the study actually do to rule that process out?

If several strong alternatives remain largely unresolved, the causal wording should remain correspondingly cautious.

07 · A Quick Checklist

Before Accepting a Causal Claim

For the proposed causal relationship, check:
Is the exposure, intervention, or proposed cause defined clearly enough to know what causal change is being claimed?
Did the proposed cause occur before the relevant outcome?
What comparison represents what would plausibly have happened without the exposure or under the alternative intervention?
How were exposed and comparison groups created, and could that process itself explain outcome differences?
What important common causes of exposure and outcome could confound the relationship?
Were important confounders measured adequately and handled according to a defensible causal rationale?
Could selection, attrition, exclusions, or missing data have distorted the causal comparison?
Were exposure and outcome measured comparably across the relevant groups and time points?
Does the analysis correspond to the design and the particular causal effect being estimated?
Were sensitivity analyses or other checks used to investigate important assumptions and alternative explanations where appropriate?
Does the causal language reflect the remaining uncertainty rather than merely the statistical significance of the estimate?
08 · Frequently Asked Questions

Questions About Causal Claims in Research

What is the difference between correlation and causation?

Correlation or association means that variables vary together in the observed data. Causation means that changing one variable would change the outcome under the relevant causal contrast. Establishing the latter requires additional design, assumptions, and evidence beyond observing an association.

Can an observational study establish causation?

Observational studies can be designed and analyzed to estimate causal effects, but causal interpretation depends on assumptions that require careful justification, including assumptions about confounding, selection, measurement, and comparability of exposure groups. Modern causal-inference methods can strengthen such analyses without making those assumptions disappear.

Does a randomized controlled trial prove causation?

A well-designed, well-conducted, and appropriately analyzed randomized trial can provide strong causal evidence for the intervention contrast studied. Randomization does not protect against every problem, however. Attrition, deviations from assigned treatment, differential measurement, selective reporting, and analytical problems can still weaken causal interpretation.

Does controlling for confounding make an observational result causal?

Not automatically. Adjustment can address measured confounders under relevant assumptions, but important confounders may be unmeasured or measured poorly, and inappropriate adjustment can introduce bias. Examine how covariates were chosen and what assumptions remain.

Does X have to happen before Y for X to cause Y?

Yes, the relevant cause must precede its effect. Establishing temporal order is therefore essential to causal interpretation. Temporal precedence alone is not sufficient because confounding, selection, measurement, and other alternative explanations may still remain.

Do the Bradford Hill criteria prove causation?

No. Hill presented his well-known considerations as viewpoints rather than hard-and-fast rules, and modern discussions caution against using them as a checklist whose completion establishes causality. They can help organize reasoning about causal evidence, particularly across a body of research, but none provides an automatic causal verdict.

What is the most important question to ask about an observational causal claim?

Ask why the exposed and comparison groups should be expected to have similar outcomes if their exposure had been the same. Then examine which assumptions, design features, measurements, and analyses support that comparability and which plausible alternative explanations remain.

Should I avoid causal language whenever a study is observational?

Not as an absolute rule. Some observational studies are explicitly designed for causal inference using frameworks and methods appropriate to that purpose. The causal language should reflect the credibility of the design and assumptions rather than follow a simple observational-equals-noncausal rule. For ordinary associational observational studies without a defensible causal design, causal wording generally requires much greater caution.

09 · The Bottom Line

A Causal Claim Requires a Credible Answer to “What Would Have Happened Otherwise?”

The Bottom Line

When a paper claims that X caused Y, do not ask only whether X and Y were associated. Ask whether the study provides a credible comparison between the observed outcome and what would plausibly have happened under the relevant alternative.

Check temporal order, treatment or exposure assignment, confounding, selection, measurement, missing data, analytical choices, and remaining assumptions. Randomization can strengthen that comparison substantially, while observational causal inference requires particularly careful justification of the assumptions connecting the observed data to the counterfactual effect. The question is not whether every alternative explanation is imaginable; it is whether the study has addressed the important alternatives well enough for its causal language to be proportionate to the evidence.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes