03 · What You Need to Know
Study Disagreement Can Be Scientifically Informative
Different Samples Produce Different Results
Most studies observe samples rather than every possible case relevant to a research question. Even when samples are drawn appropriately, they will not contain exactly the same individuals or observations.
Suppose an intervention has a modest average effect. One study may happen to include participants who respond particularly well, while another includes more participants who respond weakly. The resulting effect estimates can differ even if both studies are unbiased representations of an underlying variable process.
Larger appropriate samples can often reduce sampling uncertainty, but they do not make different studies numerically identical.
The Underlying Phenomenon May Genuinely Vary
Sometimes researchers implicitly assume that there is one fixed effect waiting for every study to estimate. Many phenomena are not that simple.
An educational intervention may work differently for younger and older learners. A treatment may have different effects depending on baseline health. A workplace policy may produce different outcomes in organizations with different cultures. A social behavior may change across historical periods.
In these cases, different results may reveal heterogeneity: genuine variation in effects, relationships, mechanisms, or experiences across populations or conditions.
The National Academies identifies inherent but previously uncharacterized variability as one scientifically informative source of non-replicability. A discrepancy can therefore expose something important about how the system behaves rather than merely indicating that one study failed.
Context Can Change the Answer
A research conclusion is always produced under some set of conditions.
Studies may take place in different countries, schools, hospitals, laboratories, organizations, communities, economic environments, or historical periods. Even when researchers ask what appears to be the same question, contextual conditions can modify the phenomenon.
This is particularly consequential in social, behavioral, educational, ecological, and health research, where outcomes often emerge from interactions among individuals and their environments.
If an intervention works in one setting but not another, the scientifically interesting question may become: Under what conditions does it work?
Studies May Not Be Testing Exactly the Same Thing
Two papers can use the same conceptual label while operationalizing it differently.
“Student engagement” might be measured through attendance in one study, self-reported engagement in another, behavioral activity on a digital platform in a third, and a multidimensional validated scale in a fourth.
Likewise, two interventions carrying the same name may differ in duration, intensity, instructor training, delivery format, supporting materials, or participant adherence.
Different results may therefore reflect substantive differences hidden beneath apparently identical terminology.
Watch Out
Before calling two studies contradictory, verify that their research questions, constructs, interventions, outcomes, populations, time frames, and comparisons are sufficiently similar for the results to be expected to agree.
Measurement Differences Can Produce Different Findings
Measurement determines what aspect of a phenomenon becomes visible in the data.
One educational study may measure immediate recall after an intervention. Another may measure retention six months later. A health study may examine a biomarker while another evaluates patient-reported quality of life. Both may be investigating the same broad intervention while assessing different outcomes.
The studies could therefore produce apparently conflicting conclusions that are actually answers to different questions.
Even when researchers intend to measure the same construct, instruments can differ in reliability, sensitivity, thresholds, calibration, or validity. The National Academies identifies limitations in measurement precision as one constraint on the replicability of scientific results.
Small Methodological Differences Can Matter
Replication does not necessarily mean that every procedural detail is perfectly identical. Indeed, some scientific questions require testing whether a finding persists when procedures change.
But seemingly minor differences can sometimes influence results. The order of survey questions, laboratory temperature, timing of measurements, instructions given to participants, software settings, inclusion criteria, researcher interactions, or implementation fidelity may alter what is observed.
The importance of such differences may not be known until studies disagree.
This is one reason a failed replication can generate new knowledge. It may reveal that an effect depends on a condition researchers had previously assumed was irrelevant.
Analytical Decisions Can Affect Results
Researchers also make decisions about how data will be processed and analyzed.
These may concern inclusion and exclusion criteria, treatment of missing observations, construction of variables, statistical models, covariates, transformations, coding procedures, qualitative interpretation, or sensitivity analyses.
Different defensible analytical approaches can occasionally produce different estimates or emphasize different aspects of complex data.
Transparency is therefore important. When analytical procedures are documented clearly, researchers can investigate whether disagreement reflects the underlying evidence or differences in how that evidence was analyzed.
Random Variation Means Exact Agreement Is Usually Unrealistic
Even if two studies were extremely similar, their results would not normally be expected to match perfectly when they use new observations.
Sampling variability and other random processes create differences between estimates. The question is therefore not whether two studies produce identical numbers, but whether their results are sufficiently consistent given the uncertainty associated with each.
The National Academies explicitly cautions against treating replication as a simple binary pass-or-fail outcome. The degree of consistency between results can be more informative than asking whether one crossed a particular statistical threshold and the other did not.
“Significant” and “Not Significant” Do Not Necessarily Mean Opposite Results
This deserves particular attention.
Suppose Study A estimates an effect of 5.2 units and reports a statistically significant result. Study B estimates an effect of 4.8 units but has greater uncertainty and does not cross the chosen significance threshold.
It would be misleading to summarize the studies as “Study A found an effect, while Study B found no effect.” Their point estimates are actually very similar.
Comparing statistical significance labels is not the same as comparing effect estimates and their uncertainty.
One or Both Studies Can Still Have Problems
Not every disagreement is scientifically interesting heterogeneity. Some discrepancies do arise from methodological weaknesses.
Problems can include biased sampling, inadequate controls, measurement error, inappropriate analysis, selective reporting, insufficient documentation, data-processing mistakes, or, less commonly, misconduct.
Scientific rigor helps reduce these possibilities. NIH guidance emphasizes robust and unbiased design, methodology, analysis, interpretation, and reporting, as well as appropriate controls and attention to relevant variables.
The key is not to assume the explanation before examining the evidence.
A Single Failed Replication Does Not Settle the Matter
The National Academies makes an important distinction: successful replication does not guarantee that an original result is correct, and a single failed replication does not conclusively refute the original claim.
If two studies disagree, researchers need to investigate possible explanations. Differences in samples, methods, measurements, contextual conditions, uncertainty, and study quality all matter.
This helps explain why good research can sometimes produce an incorrect conclusion without implying that every conflicting result must be caused by error.
The Broader Body of Evidence Matters More Than a Duel Between Two Papers
When two studies disagree, it is tempting to treat them as opponents and choose a winner. Scientific knowledge rarely develops so neatly.
The National Academies argues that the robustness of scientific knowledge is better represented by a broader web of evidence developed through multiple lines of inquiry than by replication between only two individual studies.
Researchers may therefore examine additional studies, methodological differences, research syntheses, meta-analyses where appropriate, and evidence produced through different approaches.
This is why one research study is rarely enough to provide a definitive answer, and why two conflicting studies are rarely the end of the discussion either.