01 · The Question
Is Exact Replication Stronger Than Conceptual Replication?
Suppose you want to replicate an existing study. Should you reproduce its procedures as closely as possible, or deliberately change the methods, measures, population, or setting while testing the same underlying idea?
The first strategy is commonly called direct replication and sometimes exact replication. The second is commonly called conceptual replication. They are often discussed as competing approaches, but they are better understood as tools for answering different questions.
A close replication can provide a clearer test of whether a particular result appears again under similar conditions. A conceptual replication can test whether the underlying claim survives meaningful changes in how it is investigated. Neither approach is automatically the stronger contribution.
03 · What You Need to Know
What Direct and Conceptual Replications Actually Test
“Exact replication” is an imperfect label
The phrase “exact replication” suggests that a study can be reproduced without any differences. In practice, that is rarely possible. A replication occurs at another time and normally involves new observations, participants, researchers, or circumstances. Even carefully standardized laboratory procedures cannot reproduce every feature of an earlier study.
For this reason, “direct replication” is often the more useful term. The National Institutes of Health defines direct replication for its Replication Prize as repetition of an experimental procedure in which changes are deliberate and most features are kept as consistent as possible. It defines conceptual replication as verification of the underlying hypothesis using a different method or experimental setup.
Direct replication
Repeats the original procedure as closely as practicable so that the new result can be compared with the original finding under substantially similar conditions.
Conceptual replication
Tests the same underlying hypothesis or theoretical claim while deliberately changing important aspects of how it is operationalized or investigated.
These definitions are useful, but they are not universal. A multidisciplinary review of replication research identified multiple definitions and typologies across fields, including replication as repetition, extension, and theory testing. Researchers should therefore explain what they mean rather than assume every reader uses the labels identically.
Direct replication asks whether the specific result holds up
Imagine that an experiment reports an effect using a particular intervention, outcome measure, participant population, and procedure. A direct replication attempts to preserve the features considered relevant to producing that result while collecting new data.
The purpose is not to create a perfectly identical study. It is to make the comparison sufficiently close that the replication result is informative about the original result.
The National Academies describes the most direct assessment of replicability as conducting a study following the original methods and comparing the new results with the original ones. It also emphasizes that replication is one way researchers build confidence in scientific results.
Conceptual replication asks whether the underlying claim survives a different test
A conceptual replication deliberately introduces more substantial changes. Researchers might use a different operationalization of the independent variable, another outcome measure, an alternative experimental procedure, or another theoretically appropriate way of testing the same hypothesis.
If different methods converge on the same underlying conclusion, the evidence may suggest that the finding is not merely an artifact of one particular operationalization or procedure.
NIH describes conceptual replication as verifying the underlying hypothesis of an earlier experiment using a different method or experimental setup. This makes conceptual replication particularly relevant when the scientific question concerns the underlying hypothesis rather than literal repetition of a particular procedure.
| Question |
Direct replication |
Conceptual replication |
| Primary aim |
Test a previous result under closely similar conditions |
Test the underlying hypothesis using meaningfully different methods or conditions |
| Similarity to original procedures |
Kept as high as practicable for relevant features |
Important features are deliberately changed |
| Main strength |
Provides a relatively direct comparison with the original result |
Can test whether the claim survives alternative operationalizations or implementations |
| Main interpretive challenge |
Determining which unavoidable differences could matter |
Determining whether disagreement reflects the theory or the changed methods |
| Especially useful when |
The reliability of a particular result is uncertain or important |
The robustness or theoretical scope of a claim is the central question |
Why direct replication can provide a cleaner test of a particular finding
Suppose a conceptual replication changes the manipulation, outcome measure, and participant population simultaneously and then obtains a different result. What failed to replicate?
Perhaps the original finding was unreliable. But perhaps the alternative manipulation did not engage the same process, the new outcome measured something different, or the effect genuinely depends on the population. Multiple changes create multiple explanations for disagreement.
A well-designed direct replication reduces some of that ambiguity by preserving features of the original design. This does not make interpretation automatic, but it can make the new evidence more diagnostic of whether the particular original finding is reproducible under similar conditions.
Why conceptual replication can provide evidence a direct replication cannot
A finding that appears repeatedly under one highly specific procedure may be reliable under that procedure while remaining narrow in scope. Researchers may also want to know whether the proposed explanation survives when the concept is measured or manipulated differently.
Conceptual replication can help address that question. If a theoretically predicted effect appears across appropriately different methods, the evidence may support a broader claim than repeated use of a single procedure would support.
This is one reason programs of research can benefit from both approaches. Closely matched studies can establish whether a result is reproducible under similar conditions, while deliberately varied studies can investigate robustness, mechanisms, and boundary conditions.
Conceptual replication creates an important interpretive problem
The flexibility that makes conceptual replication useful also makes it harder to interpret in some circumstances. If a substantially changed study produces a different result, researchers must decide whether the evidence challenges the original claim or merely shows that the new implementation was not equivalent in a theoretically relevant way.
This issue has produced genuine methodological debate. Nosek and Errington argue that replication should be defined by whether a study's outcome would change confidence in the claim from the original study, rather than by a simple direct-versus-conceptual distinction. They caution that some studies called conceptual replications may actually be tests of generalizability rather than replications of the original claim.
That debate is a reason to avoid treating the labels as self-explanatory. Specify the claim being tested and explain how either possible result would affect your interpretation of that claim.
Changing the population may test generalizability rather than simple replication
Researchers sometimes call a study a conceptual replication because it repeats an earlier investigation in another country, age group, institution, species, or social setting. That can be useful research, but the central question may now be whether the finding generalizes.
The National Academies distinguishes replicability from generalizability: replicability concerns consistency across studies addressing the same scientific question with their own data, whereas generalizability concerns whether findings apply in different contexts or populations.
If your principal change is the population, it may therefore be more precise to explain that you are testing whether an existing finding applies to a new population rather than relying solely on the label “conceptual replication.”
The strongest contribution may come from a sequence of replications
A single study cannot efficiently answer every question about reliability, robustness, mechanism, and generalizability. Trying to make one replication do all of these things can produce a design in which too many features change at once.
A research program can instead move from tighter tests to broader ones. A close replication can first establish whether the original effect appears again under similar conditions. Subsequent studies can deliberately vary theoretically important features to examine robustness or boundary conditions.
This sequence is not mandatory, and the appropriate order depends on the evidence already available. In research on psychological interventions, for example, methodological recommendations have proposed using direct replication to test treatment effects and conceptual replication to investigate generalizability and potential moderators.
Neither approach is automatically more original
It is tempting to regard conceptual replication as more original because it contains more obvious changes. That confuses novelty with contribution.
A direct replication can make a substantial contribution when independent evidence about an influential or uncertain finding is badly needed. A conceptual replication can make a weak contribution if its modifications are arbitrary or make the result impossible to interpret.
The relevant question is therefore not which design looks newer. As explained when considering whether replication can constitute original research, the contribution comes from the new evidence and what it allows researchers to learn.
Watch Out
Do not assume that changing more features makes a replication scientifically stronger. Every change may help test robustness or generalizability, but it may also introduce another explanation for why the new result differs from the original. Change what the research question requires, and make those changes explicit.
07 · A Quick Checklist
Before Choosing a Replication Design
Before choosing direct or conceptual replication, check:
State the precise finding, hypothesis, or theoretical claim you want the replication to evaluate.
Review the wider evidence to determine whether close independent replications already exist.
Identify which features of the original study must remain similar for the comparison to answer your question.
List every important feature you intend to change and give a scientific reason for changing it.
Ask what a consistent result would allow you to conclude and what an inconsistent result would allow you to conclude.
Distinguish a test of the original finding from a test of generalizability to a different population or context.
Avoid adding methodological differences solely to make a direct replication appear more novel.
Describe the replication procedures and deviations transparently enough that readers can judge the similarity between studies.