03 · What You Need to Know
Before comparing conclusions, reconstruct the question behind each study
Similar topics are not necessarily the same research question
Studies are often grouped together because they share a recognizable topic: online learning, exercise and depression, artificial intelligence in education, a medication and disease outcome, or social media and well-being. Topic similarity, however, does not establish question equivalence.
A research question specifies much more than a topic. In intervention research, one common framework is PICO: population, intervention, comparator, and outcome. Cochrane uses these elements to define review questions and, importantly, distinguishes the broad question of a review from the more specific questions addressed in individual syntheses and individual included studies.
That distinction is useful well beyond systematic reviews. If two studies differ materially in whom they studied, what was done or observed, what it was compared with, or what outcome was assessed, their findings may not be directly competing estimates.
Same broad topic
Two studies investigate related phenomena but may differ in the specific question, population, exposure or intervention, comparison, outcome, context, or time frame.
Comparable research question
The studies estimate sufficiently similar relationships or effects that differences between their findings can reasonably be interpreted as disagreement about the same substantive claim.
Start with the population: about whom is each claim being made?
Suppose one study reports that a digital learning intervention improves achievement among first-year university students, while another finds no meaningful effect among postgraduate students. It would be premature to write that the second study contradicts the first.
The effect might differ by prior knowledge, age, educational level, motivation, baseline performance, institutional environment, or some other characteristic. The first study may provide a credible estimate for one population and the second for another.
This is not a minor technicality. Claims about effects are often conditional on the populations from which evidence was obtained. When populations differ, ask whether those differences are merely descriptive or whether they could plausibly modify the phenomenon under investigation.
If population characteristics appear important, examine more closely whether different populations could explain why the studies disagree.
Check whether the intervention or exposure is genuinely the same
Two studies may use the same label for interventions that differ substantially in content, intensity, duration, implementation, fidelity, or accompanying support.
"Online learning," for example, can describe anything from self-paced instructional materials to synchronous classes with intensive instructor interaction. Likewise, a behavioral intervention delivered for two weeks may not be equivalent to one delivered for six months, even if both carry the same program name.
The same issue arises in observational research. Broad exposures such as screen time, physical activity, social media use, or socioeconomic status can be operationalized differently across studies.
Before declaring contradiction, inspect what participants were actually exposed to rather than relying on the terminology used in the title or abstract.
The comparator can change the question completely
Researchers sometimes focus on the intervention and overlook what it was compared against.
Consider two hypothetical trials. One asks whether a new teaching approach produces better outcomes than no structured intervention. Another asks whether that same approach performs better than an established, well-designed teaching method. A positive result in the first and little difference in the second are entirely compatible.
The combined interpretation might be that the new approach is better than receiving no structured intervention but offers little advantage over an effective alternative.
Those are different comparisons and therefore different causal questions.
Studies can use the same outcome label while measuring different things
Outcome names can create an illusion of comparability. "Achievement," "engagement," "depression," "quality of life," "productivity," or "research impact" can represent quite different operational measures.
One educational study might measure achievement using a standardized examination, while another uses course grades. One study of engagement might rely on self-reported interest, while another counts interactions recorded in a learning platform.
Even when the underlying construct is similar, differences in instruments, thresholds, scoring, reliability, and timing can affect what the study estimates.
When findings seem contradictory, inspect whether different measures could be producing different conclusions before assuming that the underlying phenomenon changed.
Time changes the meaning of an outcome
An intervention can have an immediate effect that weakens later, or little immediate effect followed by a delayed benefit. A treatment might reduce symptoms after four weeks without changing long-term recurrence. An educational intervention might improve performance on an immediate assessment without producing better retention months later.
Therefore, "improved the outcome" and "did not improve the outcome" can coexist when the outcome was measured at different time points.
When reading apparently conflicting studies, record not only what was measured but when it was measured.
Context and setting can be part of the research question
Effects may depend on where and under what conditions an intervention or exposure occurs. An educational program implemented with extensive instructor training may perform differently when introduced without comparable support. A public health intervention can operate differently across health systems. Organizational practices may have different consequences in small firms and large institutions.
Context should not become a convenient explanation invented whenever results disagree. But neither should it be ignored when there are plausible mechanisms through which contextual conditions could influence the effect.
Cochrane's guidance on complex interventions similarly recognizes that populations, interventions, outcomes, and their effects can vary across contexts, making it necessary to consider whether evidence from different settings can meaningfully address the same synthesis question.
Different research designs may be estimating different quantities
Imagine that an observational study reports a strong association between participation in a program and academic achievement, while a randomized experiment reports a much smaller effect.
The difference may reflect bias in one or both studies, but there is another issue to examine first: what exactly does each design permit researchers to estimate?
In the observational study, students who participate may differ systematically from those who do not. The randomized study, if appropriately conducted, is designed to estimate an intervention effect under its particular experimental conditions. Treating the two numerical results as interchangeable estimates can therefore be misleading.
Before interpreting the discrepancy as direct contradiction, consider how different research designs could produce apparently conflicting findings.
Statistical significance does not define contradiction
One of the easiest ways to manufacture an apparent contradiction is to compare significance labels.
Suppose Study A reports an estimated effect of 0.22 and labels it statistically significant. Study B reports an estimated effect of 0.18 but labels it nonsignificant because its estimate is less precise. It would be incorrect to conclude from those labels alone that Study A found an effect while Study B found no effect.
Altman and Bland emphasize this point in the context of comparing estimates: a statistically significant result in one analysis and a nonsignificant result in another does not itself demonstrate that the two effects differ. The relevant comparison concerns the estimates themselves and the uncertainty around their difference.
Watch Out
"Significant here but not significant there" is not a statistical demonstration of contradiction. Compare effect estimates directly and consider their uncertainty rather than treating separate p-values as competing verdicts.
Direction alone can also oversimplify the comparison
Suppose one study estimates a small positive effect and another estimates a small negative effect. The opposite signs can look dramatic on a table, particularly if the estimates are displayed simply as "+" and "−."
But if both estimates are highly uncertain and compatible with effects close to zero, the data may provide little evidence of a meaningful difference between them.
Conversely, two estimates pointing in the same direction can still differ substantially in magnitude. A modest benefit and an extremely large benefit do not become equivalent merely because both are positive.
Compare magnitude, direction, precision, and the substantive meaning of the estimates together.
A genuine contradiction concerns incompatible claims, not merely different numbers
No two independent studies will produce exactly the same numerical estimate. Sampling variability alone ensures some variation. A genuine contradiction therefore cannot mean simply that the reported numbers differ.
The stronger question is whether the studies support materially incompatible interpretations of the same underlying claim.
For example, suppose two highly comparable, well-conducted studies estimate the same intervention effect in similar populations using the same outcome and time frame. One provides reasonably precise evidence of a meaningful benefit, while the other provides reasonably precise evidence of meaningful harm. That deserves to be treated as substantive disagreement and investigated carefully.
By contrast, if one study estimates a modest benefit with substantial uncertainty and another estimates approximately no effect with substantial uncertainty, the evidence may be inconclusive rather than genuinely contradictory.
Comparability is a matter of degree
Studies do not fall neatly into categories of "identical" or "completely different." Almost every pair of studies differs somewhere.
The question is whether those differences matter for the inference you want to make. Two studies can use slightly different populations yet remain sufficiently comparable for a useful synthesis. Alternatively, one apparently modest difference, such as the comparator or outcome definition, can change the research question substantially.
Cochrane's guidance on synthesis reflects this judgment explicitly: study characteristics should be examined and compared to determine which studies are sufficiently similar to be grouped for a particular synthesis. Broader grouping can sometimes reveal meaningful variation, but inappropriate grouping can produce an average that is difficult to interpret.
There is no universal checklist that mechanically decides comparability. Substantive knowledge, methodological reasoning, and a clear definition of the claim being evaluated remain necessary.
06 · What This Means for You
How to test whether an apparent contradiction is real
When two studies appear to conflict, temporarily ignore the conclusions written by their authors. Reconstruct what each study actually asked and what its data allow it to claim.
A simple decision framework
If the populations differ substantially
Ask whether the effect could reasonably vary between those populations before describing the findings as contradictory.
If the intervention, exposure, or comparator differs
Determine whether the studies are estimating the same contrast. If not, state the conclusions separately.
If the outcomes or measurement times differ
Determine whether the outcomes represent the same construct and whether effects could reasonably change over time.
If research designs differ
Identify what each design estimates and which sources of bias could influence the comparison.
If only the significance labels differ
Do not infer contradiction. Compare the effect estimates and their uncertainty directly.
If sufficiently comparable studies provide materially incompatible estimates
Treat the disagreement as potentially genuine and investigate sampling variation, risk of bias, analytical differences, effect modification, and other plausible explanations.
Create a comparison table before writing your synthesis
A simple study-characteristics table can prevent surprisingly large interpretive errors. For each study, record the population, intervention or exposure, comparator, outcome definition, measurement time, setting, design, effect measure, estimate, uncertainty, and major methodological limitations.
Then compare the studies horizontally rather than reading them one paper at a time. Differences that were easy to overlook in prose often become obvious when placed side by side. This approach is consistent with systematic-review practice, where tabulating study characteristics helps researchers assess whether studies are sufficiently similar to contribute to the same synthesis.
Separate three possible conclusions
After comparison, you may find that the studies are addressing different questions and therefore do not directly contradict one another. You may find that they address a common question but produce different estimates that remain reasonably compatible given their uncertainty. Or you may find that sufficiently comparable studies provide materially different evidence that genuinely requires explanation.
Only the last situation should lead quickly to language such as "contradictory findings." Even then, the analysis is not finished. You still need to determine what could explain why the studies reached different conclusions.
Let disagreement make the research question more precise
Apparent contradictions sometimes reveal that the original question was too broad.
"Does this intervention work?" may become "For whom does it work, compared with what, for which outcome, over what period, and under what implementation conditions?"
That is not merely a semantic refinement. It can transform an apparently chaotic literature into a set of more specific, testable claims. Occasionally the reward for reading contradictory papers is discovering that your research question needed better manners all along.