03 · What You Need to Know
Replication Is Supposed to Test Claims, Not Produce a Predetermined Answer
The Phrase "Failed Replication" Can Be Misleading
Researchers commonly describe a replication as successful when it obtains a result sufficiently consistent with the original finding and failed when it does not. The terminology is convenient, but it can encourage the wrong mental model.
Nosek and Errington propose defining replication according to whether possible outcomes would provide diagnostic evidence about a claim from prior research. Under this view, replication is valuable because it confronts existing understanding with new data.
The important question is therefore not:
Did the replication produce the answer we wanted?
It is:
What does the new evidence do to our confidence in the claim?
An editorial in Nature Human Behaviour makes a related argument: rigorous replications contribute to scientific knowledge regardless of whether they confirm the original result. This does not mean that every study labeled a replication is automatically useful. The design still needs to be capable of producing informative evidence. But outcome direction should not determine whether the research itself is regarded as worthwhile.
Replication result
The empirical evidence produced by the new study and how it compares with the evidence supporting the prior claim.
Quality of the replication
Whether the design, implementation, measurement, analysis, and reporting provide a credible and informative test of that claim.
These are different judgments. An excellent replication can produce evidence inconsistent with the original finding. A poorly designed replication can produce the same result as the original and still tell us relatively little.
An Inconsistent Replication Usually Changes Confidence, Not Truth From True to False
Suppose an original study reports a positive effect and your replication does not obtain comparable evidence.
It is tempting to conclude:
"The original finding was false."
That conclusion is usually too strong.
The original study produced one set of evidence under particular conditions. The replication produces another. The two results must now be interpreted together. An inconsistent replication may reduce confidence in the original claim or suggest that its reliability is more constrained than previously assumed. It does not ordinarily establish which individual study represents an absolute truth.
Nosek and Errington make this point explicitly: an unsuccessful replication can indicate that the reliability of a finding is more limited than previously recognized. That is scientifically useful information.
Do Not Define Replication Failure as "Significant Before, Nonsignificant Now"
One of the weakest ways to evaluate replication is to compare statistical significance labels.
Suppose the original study reports an effect estimate of 0.30 with a p-value below a conventional significance threshold. Your replication estimates an effect of 0.27 but produces a p-value slightly above that threshold.
Declaring that the original study "worked" and the replication "failed" would ignore how similar the effect estimates might actually be.
The reverse can also happen. Two studies may both produce statistically significant results while estimating substantially different effect magnitudes.
Large replication projects illustrate why richer criteria are needed. Camerer and colleagues, for example, evaluated replications of social science experiments using multiple indicators rather than relying on one significance comparison. Their replication effect sizes were, on average, smaller than the original estimates.
Replication assessment should therefore consider effect estimates, uncertainty, direction, compatibility between studies, the prespecified replication criterion, and the substantive meaning of the difference.
Watch Out
"The original result was significant but mine was not" does not by itself demonstrate that the studies produced meaningfully different effects. Compare the estimates and their uncertainty, and use an appropriate prespecified criterion for evaluating replication.
A Smaller Effect Can Be an Important Replication Result
A replication does not need to estimate an effect of zero to challenge the interpretation of an original study.
Suppose the original research reports a very large effect that makes an intervention appear transformative. A larger replication estimates the effect in the same direction but substantially smaller.
Has the study replicated?
The answer depends on the criterion and the claim being evaluated. If the claim was merely that the effect is positive, the replication may support it. If the practical argument depended on the effect being large, the new evidence may materially weaken that interpretation.
This is why replication should not be reduced to a binary badge. Effect magnitude matters, particularly when findings are used to justify interventions, costs, policies, or theoretical claims about the strength of a phenomenon.
An Inconsistent Result Can Reveal a Boundary Condition
Imagine that an effect appears repeatedly among novice learners but not in your replication involving experts.
One interpretation is that the original claim does not generalize as broadly as expected. Expertise may be a boundary condition.
Likewise, a finding might appear in one educational system but not another, with one implementation but not another, or under one measurement approach but not another.
Such discrepancies can generate useful theoretical questions:
- Which condition differs between the studies?
- Could that condition plausibly influence the phenomenon?
- Was the original claim stated too broadly?
- What additional study could distinguish competing explanations?
A non-replication can therefore transform an apparently simple claim into a more precise one. Instead of "X affects Y," the emerging evidence may suggest "X affects Y under these conditions but not necessarily under those conditions."
This is particularly informative when the replication deliberately changes a population or context to test whether the original claim generalizes beyond the population first studied.
But Do Not Invent a Boundary Condition After Every Inconsistent Result
There is a danger on the other side.
Once a replication differs from the original result, researchers can inspect every difference between the studies and construct a plausible explanation for why the effect "should" have appeared in one but not the other.
Perhaps the participants were slightly older. Perhaps the classroom was different. Perhaps the study occurred in another country. Perhaps the software version changed. Given enough methodological details, some post hoc story can almost always be told.
That does not make the explanation correct.
A plausible moderator or boundary condition generated after observing the discrepancy should usually be treated as a hypothesis requiring further evidence. Ideally, theoretically important differences are identified before the replication and their predicted consequences are specified in advance.
The discrepancy generates the question. It does not automatically answer it.
Sampling Variation Remains a Possible Explanation
Two well-conducted studies sampling from the same underlying process will not produce identical estimates.
Some disagreement between an original study and a replication can therefore occur because of sampling variation alone. The smaller and less precise the studies, the wider the range of estimates that may reasonably occur.
This is another reason to examine confidence intervals or other appropriate measures of uncertainty rather than comparing point estimates mechanically.
If the original study estimated a large effect with considerable uncertainty and the replication estimated a smaller effect with considerable uncertainty, the studies may be less contradictory than their headlines suggest.
Conversely, precise estimates that are strongly incompatible may provide much stronger evidence that something differs between the studies or that the original claim requires revision.
Methodological Differences Can Matter, but They Need Symmetrical Treatment
Every replication differs from the original study in some respects. Participants change. Time passes. Researchers differ. Materials or technologies may require adaptation.
Those differences can matter when interpreting an inconsistent result.
But methodological differences should be considered symmetrically.
If you would have regarded a result consistent with the original finding as evidence supporting the claim despite those differences, you should be cautious about declaring the same differences fatal only after the replication produces an inconsistent result.
This is why important departures should be documented and their expected implications considered before data collection whenever possible.
If substantial adaptations were unavoidable, interpretation should reflect how those differences affect the evidential relationship between the studies.
A Failed Replication Does Not Establish Research Misconduct
An inconsistent replication result is evidence about a scientific claim. It is not, by itself, evidence that the original researchers fabricated data, manipulated analyses dishonestly, or committed another form of research misconduct.
Many ordinary features of research can produce discrepant findings: sampling variability, limited statistical precision, measurement differences, contextual variation, analytical flexibility, implementation differences, publication processes, or an original effect that is less robust than initially believed.
Research misconduct is a separate allegation requiring appropriate evidence and investigation.
Watch Out
Do not use a non-replication as evidence of misconduct unless independent evidence actually supports that allegation. Replication evaluates claims; it is not a misconduct adjudication procedure.
The Replication Itself Can Also Be Wrong
Replication is not epistemically privileged merely because it comes second.
The new study can contain measurement problems, implementation failures, coding errors, analytical mistakes, inadequate statistical precision, or other limitations. A later study does not automatically override an earlier one.
This matters because discussions of replication can accidentally create an asymmetry: the original result must prove itself, while the replication is treated as the final judge.
A better approach evaluates both studies critically.
Were the procedures implemented as intended? Were measures appropriate? Was the replication sufficiently informative? Were important deviations documented? Was the analysis prespecified or otherwise well justified? Do the data and materials permit appropriate scrutiny where sharing is possible?
The scientific question concerns the accumulated evidence, not which paper was published last.
Weak Replications Can Produce Ambiguous Non-Replications
Suppose the original study reports a moderate effect, but your replication has a very small sample and produces a highly uncertain estimate compatible with both a meaningful effect and little effect.
Calling this a failed replication may exaggerate what the study establishes.
The result may simply be inconclusive.
Evidence inconsistent with the original claim
The replication provides informative evidence that lowers confidence in the prior claim or its expected magnitude.
Insufficiently informative evidence
The replication is too imprecise, compromised, or ambiguous to distinguish adequately among important possibilities.
This distinction is critical. "We did not obtain significant evidence for the effect" and "we obtained evidence that the effect is absent or substantially smaller" are not equivalent statements.
A Non-Replication Can Expose an Overly Broad Original Claim
Sometimes the problem is not that the original finding was incorrect under its original conditions. The problem is that researchers, later papers, or even the original authors generalized it too broadly.
Suppose an effect is observed among one population and is subsequently discussed as though it characterizes people generally. A rigorous replication in another relevant population does not find comparable evidence.
The new result may suggest that the original phenomenon is population-dependent.
That does not erase the first study. It changes the scope of the claim that the evidence can reasonably support.
This is one of replication's most productive functions: converting broad assumptions into claims with better-specified boundaries.
A Non-Replication Can Redirect Future Research
If an influential finding does not survive a rigorous independent test, future researchers may reconsider whether it should continue serving as a foundation for new hypotheses, interventions, or models.
The new evidence can also generate more focused questions.
Perhaps the effect depends on implementation fidelity. Perhaps it is substantially smaller than originally estimated. Perhaps a contextual moderator deserves direct testing. Perhaps the theoretical mechanism needs revision. Perhaps a measurement assumption requires closer scrutiny.
The replication has therefore not produced "nothing." It has changed the next sensible research question.
This is one reason replication can be more informative than immediately pursuing novelty when an important existing claim remains uncertain.
Large Replication Projects Show Why Outcomes Should Be Treated as Evidence
Replication projects involving multiple studies illustrate the cumulative perspective particularly well.
Camerer and colleagues conducted high-powered replications of 21 experimental social science studies published in Nature and Science. Thirteen produced a significant effect in the same direction as the original under one replication criterion, while the replication effect sizes were on average approximately half the original effect sizes.
The scientific value of the project did not lie only in identifying which individual studies met a binary replication criterion. The results also provided evidence about effect-size inflation, patterns of replicability, and the ability of researchers to anticipate replication outcomes.
This illustrates a broader principle: replication evidence can teach us about both individual claims and the research processes that produce them.
Publication Should Not Depend on Whether the Original Finding Reappears
If replication studies are considered publishable only when they contradict previous work, a new form of outcome-based publication bias can emerge. If they are publishable only when they confirm previous work, the bias simply points in the opposite direction.
Nature Human Behaviour has explicitly argued that rigorous replications should be valued regardless of whether they confirm the original findings. Registered Reports offer one mechanism for reducing outcome dependence because the study's question and methods can be peer reviewed before the results are known, with publication decisions based primarily on the importance of the question and quality of the proposed methodology rather than the eventual direction of the findings.
This is especially well suited to replication research because the scientific value of the study should be established before anyone knows whether the original result will recur.
The Most Useful Question Is Often "What Do the Studies Together Suggest?"
After an inconsistent replication, researchers can become trapped in a contest:
Original study versus replication. Which one wins?
Cumulative science asks a different question.
What do both studies, along with the rest of the relevant evidence, imply about the claim?
Perhaps the original effect was overestimated. Perhaps the replication is unusually low. Perhaps the effect varies meaningfully across settings. Perhaps the combined evidence remains inconclusive. Perhaps several later studies make the apparent conflict much easier to interpret.
Where multiple sufficiently comparable studies exist, evidence synthesis can help. A meta-analysis or other appropriate synthesis may estimate the effect and heterogeneity more informatively than choosing one study as definitive.
The replication should therefore be interpreted as part of an evolving evidence base rather than as a final appeal court for the original paper.