03 · What You Need to Know
Conflicting Results Are a Scientific Problem to Explain, Not a Vote to Settle
Research findings are expected to vary. Different studies use different samples, measurements, procedures, analytical models, settings, and periods. Even studies estimating the same underlying effect will not produce numerically identical results because of sampling variation.
The important question is therefore not whether studies differ at all. It is whether the differences are large or consequential enough to change what can reasonably be concluded.
This is closely related to why well-conducted studies can reach different conclusions. Disagreement becomes especially important when plausible explanations for the differences remain unresolved.
Not every numerical difference is a genuine conflict
Suppose one study estimates an effect of 0.20 and another estimates 0.25. Their point estimates differ, but both may be compatible with essentially the same underlying effect once their uncertainty is considered.
Likewise, one study may report a conventionally statistically significant result while another does not, even though their effect estimates are similar. A difference between “statistically significant” and “not statistically significant” does not itself establish that the studies found meaningfully different effects.
Researchers should compare effect estimates, uncertainty intervals, study characteristics, and the substantive implications of the results rather than classifying studies solely according to p-value thresholds.
Sometimes studies conflict because they are not actually asking the same question
Two papers may use similar titles while studying meaningfully different phenomena. One intervention may be more intensive. One population may have a different baseline risk. Outcomes may be measured differently. Follow-up periods may differ. Implementation conditions may not be comparable.
In such situations, the apparent contradiction may partly disappear once the research questions are specified more precisely.
Instead of asking, “Does the intervention work?” the evidence may support a more informative question: “For whom, under what implementation conditions, for which outcomes, and over what period does the intervention appear beneficial?”
Some inconsistency reflects genuine heterogeneity
In evidence synthesis, heterogeneity refers broadly to variation among studies. Clinical or substantive heterogeneity can arise from differences in participants, interventions, exposures, outcomes, or settings. Methodological heterogeneity can arise from differences in study design and conduct. Statistical heterogeneity concerns variation in effect estimates beyond what would be expected from sampling variation alone.
Cochrane guidance emphasizes that heterogeneity must be considered when interpreting a meta-analysis because meaningful between-study variation affects the extent to which a single generalizable conclusion can be formed.
This is not merely a statistical nuisance. Heterogeneity can be scientifically informative. If an intervention works in one context but not another, averaging those results without investigating why may conceal the most important finding.
Unexplained inconsistency should reduce confidence
Frameworks for evaluating bodies of evidence explicitly recognize inconsistency as a reason for caution. GRADE assesses certainty by considering domains that include risk of bias, inconsistency, indirectness, imprecision, and publication bias.
If credible studies produce substantially different estimates and researchers cannot explain those differences satisfactorily, confidence in a single summary conclusion may need to decrease.
This does not mean every inconsistent literature deserves low certainty. The importance of inconsistency depends on the magnitude and direction of the differences, their consequences for decisions, and whether credible explanations account for them.
Explained variation
Study results differ in a pattern that can reasonably be linked to meaningful differences in populations, interventions, exposures, methods, outcomes, or contexts.
Unexplained inconsistency
Credible studies produce consequentially different results, but available evidence does not yet provide a satisfactory explanation for why.
A pooled average can conceal the conflict
Meta-analysis can be useful when studies differ because it allows researchers to quantify the evidence and investigate variation. But a pooled estimate does not make heterogeneity disappear.
Imagine that some credible studies find a substantial benefit while others find little benefit or possible harm. An average effect near zero could create the impression that the intervention simply has no effect. Yet the scientifically important reality may be that effects differ meaningfully among contexts or populations.
Cochrane specifically cautions that when there is considerable variation, particularly variation in the direction of effects, quoting an average intervention effect may be misleading. In some circumstances, not performing a meta-analysis may be more appropriate than producing a summary number that obscures important differences.
Watch Out
A summary estimate is not a substitute for understanding variation. When credible studies point in meaningfully different directions, an apparently neat average can conceal the uncertainty or contextual differences that matter most.
Statistical measures of heterogeneity do not decide the issue for you
Measures such as I² can help characterize inconsistency in meta-analysis, but they should not be treated as mechanical verdicts. Cochrane cautions that simple thresholds can be misleading because the importance of heterogeneity depends on factors such as the magnitude and direction of effects and the strength of evidence for variation.
With only a few studies, estimates of heterogeneity themselves may be highly uncertain. Conversely, with many studies, statistical tests may detect small differences that have little substantive importance.
The question remains interpretive: does the observed variation materially change what researchers can conclude?
Conflict may reveal a boundary condition rather than a failed replication
Suppose an educational intervention consistently improves outcomes when instructors receive extensive implementation support but shows little benefit when introduced with minimal support.
Calling the second set of studies “failed replications” may miss the more interesting explanation. Implementation support could be a boundary condition determining when the intervention produces its effect.
Scientific progress sometimes occurs when apparent contradictions are transformed into conditional explanations. Instead of asking which study is correct, researchers identify the conditions under which each pattern appears.
Conflict can reveal that an original claim was too broad
Early evidence often encourages broad statements because researchers have not yet observed enough variation to identify important qualifications.
Later studies may show that an effect differs by age, setting, dosage, measurement approach, baseline risk, implementation quality, or another factor. The original claim may then need narrowing rather than complete rejection.
This is one way new evidence can become strong enough to change what researchers previously believed. The change may concern the scope of a conclusion rather than its complete direction.
Sometimes methodological quality does explain the conflict
Preserving uncertainty does not mean pretending all conflicting studies deserve equal weight.
If studies producing one result have a serious and well-supported source of bias while stronger studies consistently produce another, the methodological difference may justify greater confidence in one pattern.
Similarly, one group of studies may directly address the research question while another provides only indirect evidence. One result may arise primarily from an unreliable measurement process. A transparent appraisal can therefore resolve some apparent conflicts without resorting to simple study counting.
The key is that the preference should be justified by evidential differences, not selected because one conclusion is more appealing.
Conflict may appropriately increase uncertainty even when one side has more studies
Ten studies supporting one result and three supporting another do not create a ten-to-three scientific vote.
The three studies might use stronger designs, different populations, or methods that address a limitation shared by the ten. Conversely, the ten might collectively provide far more credible evidence. Numbers alone cannot tell you which interpretation is warranted.
This follows from the broader principle that more evidence does not automatically mean better evidence.
Conflict should sometimes change the question
When studies repeatedly disagree, continuing to ask which single effect is “the true effect” may become unproductive.
A better question may concern variation itself. Why do effects differ? Which study characteristics predict larger or smaller estimates? Does the intervention help one population but not another? Are different measurement approaches capturing different aspects of the phenomenon?
In this way, disagreement can generate a more mature research question. Scientific uncertainty is not merely an empty space waiting to be filled; sometimes it tells researchers where the structure of the phenomenon has not yet been understood.
| Pattern in the evidence |
Reasonable response |
Why? |
| Point estimates differ slightly but uncertainty intervals overlap substantially |
Do not assume genuine conflict |
The observed differences may be compatible with sampling variation |
| Credible studies show materially different effects for identifiable populations |
Consider a conditional conclusion |
The effect may genuinely vary across populations |
| High-risk-of-bias studies differ from stronger studies |
Weight methodological credibility explicitly |
The discrepancy may partly reflect bias |
| Credible comparable studies point in different directions with no clear explanation |
Increase uncertainty |
The evidence does not yet support one confident general conclusion |
| Different methods produce different results |
Investigate methodological explanations |
The finding may depend on measurement, design, or analytical assumptions |
| Effects vary greatly across contexts |
Investigate heterogeneity rather than relying only on the average |
A single pooled estimate may conceal meaningful variation |
Greater uncertainty can be a better scientific conclusion
Researchers sometimes feel pressure to resolve a literature into a clear positive or negative answer. Yet evidence appraisal does not require certainty where the evidence does not provide it.
If credible evidence remains inconsistent and important explanations for that inconsistency have not been resolved, increasing uncertainty is not indecision. It is an evidence-responsive conclusion.
The remaining task is then to identify what evidence would distinguish among the plausible explanations.
06 · What This Means for You
Investigate the Conflict Before Deciding What It Means
When credible studies disagree, begin by specifying exactly how they disagree. Compare effect estimates and uncertainty rather than relying only on whether authors described their findings as positive, negative, or statistically significant.
Then examine differences in populations, interventions or exposures, comparison conditions, outcomes, follow-up periods, study designs, measurements, analytical decisions, and risk of bias. Some conflicts become understandable once those differences are made explicit.
If the discrepancy remains consequential and unexplained, allow it to reduce your confidence.
A simple decision framework
If apparently conflicting studies have similar estimates but different significance labels
Compare the estimates and uncertainty directly before concluding that genuine inconsistency exists.
If methodological quality clearly differs between groups of studies
Consider whether the methodological difference plausibly explains the conflicting results.
If results differ systematically across populations or contexts
Consider a conditional conclusion and investigate effect modification or boundary conditions.
If credible comparable studies differ substantially and no convincing explanation exists
Reduce confidence in a single general conclusion and preserve the uncertainty.
If a pooled average conceals effects in opposite directions
Do not interpret the average without examining the heterogeneity and its practical implications.
Sometimes the most useful outcome of conflicting evidence is a better question. Instead of asking which result is correct, ask what conditions could make both results scientifically understandable.