03 · What You Need to Know
Why Replication Can Be a Legitimate Response to a Research Gap
Replication and Reproducibility Are Not Necessarily the Same Thing
The terminology surrounding replication and reproducibility varies across disciplines, and some research communities use the terms differently or even in opposing ways.
For its 2019 consensus report Reproducibility and Replicability in Science, the National Academies of Sciences, Engineering, and Medicine adopted a specific distinction. It defined reproducibility as obtaining consistent computational results using the same input data, computational steps, methods, code, and conditions of analysis. It defined replicability as obtaining consistent results across studies aimed at answering the same scientific question, with each study obtaining its own data.
Under that terminology, recomputing an analysis from the original data is reproducibility; conducting another study with new data to investigate the same scientific question is replication.
Because terminology differs by field, use the definitions expected in your discipline where necessary. More important than the label is explaining exactly what your proposed study repeats, what it changes, and what uncertainty it is intended to test.
One Study Rarely Eliminates Every Relevant Uncertainty
A study can be well designed and still produce an estimate with uncertainty. Its sample may differ by chance from the wider population. Measurements have limitations. Implementation can vary. Analytical decisions may matter. Conditions present in one study may not occur elsewhere.
The National Academies emphasizes that scientific communities gain confidence over time through scrutiny and repeated testing of scientific claims. Replication is one important way to build confidence in the scientific merit of a result, although it is not the only way science establishes reliable knowledge.
This creates a straightforward reason replication can address a gap: the original evidence exists, but confidence in the relevant claim remains incomplete.
The Gap Is Uncertainty, Not Absence
Suppose one study reports an important association. A second independent study has never examined the same question.
It would be inaccurate to claim that “no one has studied this before.” Someone clearly has. But there may still be an evidence gap if an important conclusion rests heavily on one study and independent evidence would materially affect confidence in it.
Missing-research gap
The relevant question has not been investigated adequately at all.
Replication-related gap
An important finding exists, but uncertainty remains about whether it can be obtained consistently with new data or under relevant conditions.
This is why a worthwhile study does not necessarily require a completely untouched topic.
Replication Can Test Whether a Result Is Reliable
The most obvious reason to replicate a finding is to determine whether another study addressing the same scientific question obtains a result consistent with the original evidence.
The National Academies notes that replication is one of the key ways scientists build confidence in results: consistency across studies makes a claim more credible as reliable scientific knowledge.
This is particularly relevant when the original finding is consequential, surprising, influential, or supported by relatively little independent evidence.
Replication does not guarantee truth. A successful replication does not prove that every interpretation of the original study is correct. But consistent evidence from independent studies can reduce some forms of uncertainty that a single result cannot resolve alone.
A Failed Replication Does Not Automatically Prove the Original Study Was Wrong
The opposite interpretation also requires caution.
The National Academies explicitly warns that a single failure to replicate does not conclusively refute an original claim. Results can differ for many reasons, including inherent variability, measurement precision, differences in methods, populations, conditions, or previously unrecognized phenomena.
A replication therefore should not be framed as a simplistic pass-or-fail examination of another researcher.
When findings differ, the scientifically useful question is often why. The discrepancy may reveal a problem with the original evidence, a problem with the replication, or a meaningful condition under which the effect changes.
Watch Out
Do not describe one successful replication as final proof or one unsuccessful replication as definitive disproof. Interpret the studies together, including their uncertainty, methods, measurements, populations, and conditions.
Replication Can Strengthen Weak Evidence
A question may already have several studies yet remain supported by insufficient evidence.
Perhaps the existing studies are small. Perhaps most were conducted by the same research group. Perhaps an important result has not been independently tested. In such situations, replication can contribute additional evidence and potentially improve confidence in the conclusion.
The important point is that another study should address the reason confidence remains limited. If the existing evidence is weak because every study uses a seriously inadequate measurement, simply repeating that measurement may do little to improve the evidence.
This connects replication to the broader question of whether weak evidence can represent a more consequential gap than completely missing evidence.
Replication Can Test Generalizability When Conditions Change
Researchers sometimes repeat a study in another population, institution, country, laboratory, or other setting. This can test more than whether the original result appears again under nearly identical conditions.
The National Academies distinguishes replicability from generalizability, which concerns the extent to which study results apply in different contexts or populations. A study using new data in a different context may therefore contain elements of both replication and a test of generalizability.
For example, if a relationship has repeatedly been observed among university students, studying it among working adults can help determine whether the finding extends to another population when that population difference is substantively relevant.
The justification should explain why the changed condition matters. Simply changing the population or country does not automatically make a replication informative.
Replication Can Reveal Boundary Conditions
When a result appears in one context but not another, the difference can expose conditions under which the underlying claim holds.
Suppose an intervention works in well-resourced organizations but produces substantially different results where staffing and infrastructure are constrained. The second study may reveal that the effect depends on implementation conditions.
Rather than concluding simply that one study “replicated” and another “failed,” researchers can investigate whether the difference identifies a boundary condition.
This makes replication especially useful for theory testing. A study can examine whether a theoretical prediction continues to hold under meaningfully different conditions.
Replication Can Test Robustness to Methodological Variation
Researchers may also investigate whether a conclusion persists when aspects of the method change.
Perhaps an original result depends on one particular measurement instrument, analytical specification, recruitment method, operational definition, or laboratory procedure. A carefully designed follow-up study can test whether the conclusion survives a defensible methodological variation.
This type of research asks a different question from an extremely close replication. Instead of asking only whether the result appears again under highly similar conditions, it asks whether the finding is robust to a change that should not eliminate the phenomenon if the underlying claim is sufficiently general.
The more the method changes, however, the harder it can become to determine why results differ. Researchers should therefore state what is intentionally varied and what inference that variation permits.
Replication Types Are Not Named Consistently Across Fields
You may encounter labels such as direct replication, exact replication, close replication, conceptual replication, or constructive replication. These terms do not have perfectly standardized meanings across disciplines.
A practical way to avoid unnecessary terminology disputes is to describe the design itself.
| Replication approach |
What remains similar |
Main uncertainty it can address |
| Close repetition |
Question, procedures, measures, and conditions are kept as similar as practical |
Whether a result can be obtained consistently with new data under similar conditions |
| Independent repetition |
The same central question is tested by another sample, team, laboratory, or dataset |
Whether the finding depends unusually on the original study or investigators |
| Methodologically varied replication |
The central claim is retained while a relevant method or operationalization changes |
Whether the finding is robust to a defensible methodological variation |
| Contextually varied replication |
The central question remains while population or setting changes |
Whether the finding generalizes or depends on contextual conditions |
The labels used for these approaches vary. The stronger proposal explains what is held constant, what is changed, and why that design answers an important uncertainty.
Replication Is Particularly Useful When a Finding Matters
Not every published result deserves an immediate replication study.
The case becomes stronger when confidence in the finding has substantial consequences. The result may influence theory, policy, clinical practice, professional decisions, subsequent research, or widely used interventions. It may also be unusually surprising or foundational to a larger literature.
If many later claims depend on an uncertain result, independent testing can have greater value than replicating a low-consequence finding simply because replication is possible.
This is the same distinction that applies to other gaps: existence and importance are separate questions.
Replication Can Be Valuable Even When Previous Replications Exist
One successful replication does not necessarily eliminate every replication-related research need.
Researchers may still be uncertain about generalizability, effect magnitude, important subgroups, robustness across methods, or performance under different conditions. Conversely, repeated successful tests across sufficiently varied and relevant conditions may eventually make another similar replication relatively low value.
There is no universal number of replications after which a claim becomes “confirmed.” The appropriate amount of evidence depends on the question, uncertainty, consequences, and characteristics of the evidence.
A Replication Should Not Be Justified Merely by Saying “More Research Is Needed”
A strong replication proposal specifies what uncertainty another study will reduce.
Compare these rationales:
| Weak rationale |
What is missing |
Stronger rationale |
| The study should be replicated. |
Why replication is needed |
An influential finding currently rests on one study and lacks independent evidence using new data. |
| More research is needed. |
What remains uncertain |
Existing estimates are imprecise, and another adequately sized study could materially improve the evidence. |
| The study has not been replicated in our country. |
Why country matters |
A specified contextual condition differs and may alter the mechanism or applicability of the finding. |
| We will use a different method. |
Why methodological variation matters |
The original conclusion depends on one operationalization, leaving uncertainty about whether the finding is robust to another valid measure. |
Replication Can Address a Gap Without Claiming Novelty
Replication exposes a weakness in the idea that every valuable study must be unprecedented.
If a consequential claim is uncertain, obtaining additional independent evidence can make an important contribution even though the research question is already known. The contribution lies in verification, robustness, precision, generalizability, or clarification of boundary conditions.
This is why the stronger question is not whether no one has studied the topic before. It is whether the current evidence is strong enough for the conclusion researchers want to draw.
Replication Does Not Have to Produce the Same Numerical Result
Studies using new data should not be expected to produce identical estimates.
The National Academies defines replicability in terms of consistent results given the uncertainty inherent in the system under study. Assessing consistency therefore requires more than checking whether two studies report the same number or whether both cross a statistical-significance threshold.
Researchers should consider effect estimates, uncertainty, design, measurements, and the scientific question being tested. What counts as sufficiently consistent can also vary among fields.
A Replication Can Create a New Research Gap
Replication sometimes resolves one uncertainty while revealing another.
If a result replicates in one setting but not another, researchers may need to investigate which contextual difference matters. If a finding appears with one measurement but not another, measurement may become the new question. If independent replications produce materially conflicting findings, inconsistency itself may become an evidence gap.
This is how replication participates in science's self-correcting process. It does not merely stamp findings as successful or failed; it can expose new questions about mechanisms, boundaries, methods, and evidence quality.
When credible replication results disagree, the next issue may be whether the conflicting findings themselves constitute a research gap.
Replication Is One Part of a Larger Evidence Base
Replication is valuable, but scientific confidence should not be reduced to whether one study has been successfully repeated once.
The National Academies emphasizes that the robustness of science is better represented by a broader web of knowledge reinforced through multiple lines of inquiry than by replication between only two individual studies. Replication is one way to gain confidence, alongside evidence synthesis, triangulation, theoretical testing, methodological scrutiny, and other forms of investigation.
The goal is therefore stronger knowledge, not replication for its own sake.