03 · What You Need to Know
Exact Replication Is Not the Standard You Actually Need
There Is No Such Thing as a Literally Exact Replication
Replication is sometimes described as repeating an original study exactly. That description is intuitive but misleading.
Nosek and Errington argue explicitly that exact replication does not exist. Every new study necessarily differs from its predecessor in some combination of participants or other units, treatments, outcomes, settings, researchers, historical circumstances, or implementation details.
Even if you use the same protocol, materials, equipment, eligibility criteria, and analytical code, you will collect new data at another time from different observations. Those differences are not defects in replication. New evidence is the point.
The more useful definition focuses on the prior claim. Under Nosek and Errington's framework, a study functions as a replication when possible outcomes would be diagnostic of that claim: an outcome consistent with the previous evidence would increase confidence in it, while an inconsistent outcome would decrease confidence in it.
Literal duplication
An impossible ideal in which every condition of the original study would somehow be recreated identically.
Informative replication
A new study that preserves or appropriately recreates the conditions needed for its possible outcomes to provide evidence about the prior claim.
This distinction changes the problem. You do not need to ask, "Did I copy everything?" You need to ask, "Do the differences prevent this study from testing what I say it tests?"
First Identify Exactly What You Are Trying to Replicate
Before deciding whether a methodological difference matters, define the replication target.
A paper may contain several experiments, hypotheses, outcomes, analyses, and theoretical claims. Saying "I am replicating this paper" is often too broad to guide methodological decisions.
Instead, write the target claim in one sentence.
For example:
Students receiving intervention X performed better on outcome Y than students receiving the comparison condition under the conditions examined.
Now methodological decisions can be evaluated against that claim. Does changing the software affect the intervention? Does translating the instrument alter measurement of Y? Does recruiting another population test the same claim or a broader one? Does changing the comparison condition alter the contrast altogether?
Without a clearly specified target, every difference can look equally important. Once the target is explicit, the hierarchy of methodological features becomes much easier to reason about.
Separate Essential Features From Incidental Features
Original studies contain innumerable details. Not all of them are equally relevant to the phenomenon.
Suppose an experiment was conducted at 10:00 a.m. in a room with white walls using desktop computers. Those features may be incidental to a claim about retrieval practice. They could be critical, however, in research concerning circadian effects, environmental perception, or device-specific interaction.
Whether a difference matters therefore depends on theory and evidence.
Nosek and Errington describe replication in terms of variation across units, treatments, outcomes, and settings. Researchers inevitably make judgments about which variations should and should not affect the claim. When theory is immature, those judgments can be uncertain, which is one reason close adherence to the original procedure may be especially useful.
| Feature |
Question to Ask |
Possible Consequence |
| Participants or units |
Does the new sample differ in a way that could affect the phenomenon? |
May alter generalizability or reveal a population boundary |
| Treatment or manipulation |
Does the adapted procedure still instantiate the same intervention or experimental contrast? |
May change what causal or substantive effect is being tested |
| Outcome or measure |
Does the alternative instrument represent the same construct adequately? |
May change the meaning or comparability of the outcome |
| Setting |
Could the new environment plausibly influence the effect? |
May be incidental or may test a boundary condition |
| Analysis |
Does the revised analysis estimate the same target quantity? |
May alter the inferential question even with identical data |
This is not a universal checklist of what must remain identical. It is a framework for asking why a difference matters.
Try to Recover the Original Materials Before Reconstructing Them
If something appears unavailable, do not immediately invent a replacement.
Check the article's supplementary files, journal page, repository links, preregistration, study protocol, appendices, data and materials availability statements, project websites, and related publications. Search repositories using the study title, authors, project name, or persistent identifiers.
If important materials or procedural details remain unavailable, contacting the corresponding author or another member of the original research team may be appropriate. Brandt and colleagues' Replication Recipe recommends communicating with original authors as part of planning a convincing close replication, particularly to obtain original materials and clarify methodological details.
Author contact should not be treated as permission to conduct the replication. The purpose is methodological clarification.
If you receive additional materials or information, document what was obtained and how it informed the replication. If the authors do not respond or materials no longer exist, document that too. A transparent unresolved uncertainty is preferable to silently guessing what the original researchers probably did.
Missing Methodological Details Require Reconstruction, Not Pretend Certainty
Sometimes the problem is not that a material is unavailable but that the article never tells you enough to reproduce the procedure.
For example, the paper might state that participants received "standardized instructions" without providing them. An intervention may be described only in broad terms. Exclusion criteria may be unclear. The precise sequence of tasks may be omitted.
When those details cannot be recovered, you must make methodological decisions.
Use the available evidence to reconstruct the procedure as defensibly as possible. Distinguish clearly between features documented by the original source and features you had to infer or create. Where several reasonable interpretations exist, explain which one you chose and why.
Watch Out
Do not fill gaps in the original method and then describe the result as an exact replication. Unreported details are themselves a source of uncertainty. Record your reconstruction decisions so readers can judge whether they could plausibly explain differences between the studies.
Outdated Technology May Need Functional Rather Than Literal Reproduction
Replication becomes particularly interesting when the original study depends on technology that is obsolete, unavailable, unsupported, or incompatible with current systems.
Suppose an experiment used software that no longer runs on modern computers. You might recreate the task using another platform.
The important question is what must remain functionally equivalent. If timing is critical, does the new platform reproduce the relevant timing accurately? If visual presentation matters, are stimuli displayed comparably? If participant interaction is central, does the replacement interface change how participants complete the task?
Simply saying that the new software "does the same thing" is not enough when implementation details could affect the outcome.
Document the replacement and, where possible, validate relevant features before conducting the main study.
Translation Is an Adaptation, Not Just a Change of Words
Suppose the original study was conducted in English and your participants require another language. Administering the original English materials merely to preserve literal similarity could undermine comprehension and therefore make the replication less informative.
Nosek and Errington use precisely this kind of example to illustrate why procedural identity cannot define replication. A study conducted with Tagalog-speaking participants in the Philippines may require translation if the research claim concerns something other than English-language comprehension.
But translation can change meaning, difficulty, tone, cultural assumptions, or psychometric properties. It should therefore be treated as a substantive methodological adaptation rather than clerical substitution.
Use an appropriate translation and adaptation process for the type of material involved. If a validated version already exists, determine whether it is suitable for your population and purpose. For measurement instruments, consider whether scores remain interpretable as representing the intended construct rather than assuming linguistic translation guarantees measurement equivalence.
A Different Population May Still Provide Replication Evidence
You may be unable to access the original population. Perhaps the first study involved students at a particular institution, participants from another country, members of a specialized profession, or a historical cohort that can no longer be recruited.
A different population does not automatically prevent replication.
If the prior claim is expected to hold in the new population and your study provides a meaningful test of that expectation, the result can still provide replication evidence. The population difference may simultaneously test some degree of generalizability.
The key is to explain why the difference should or should not matter. A population change that is theoretically irrelevant poses a different interpretive problem from one that deliberately crosses a suspected boundary condition.
If population differences are central to your design, consider more carefully whether the new population still tests the original claim or creates a different research question.
Ethical Standards Override Methodological Fidelity
Some original procedures should not be reproduced.
Ethical standards, institutional requirements, professional norms, or legal rules may have changed since the original study. A procedure considered acceptable when the study was first conducted may no longer be permissible, or your research ethics committee may require additional protections.
You should not compromise participant protection simply to maximize methodological similarity.
Instead, identify the required change and assess its implications for the replication. If the ethical modification affects an element plausibly central to the phenomenon, acknowledge that the new study is not a close recreation of that feature and interpret discrepancies accordingly.
Ethical constraints can therefore limit how directly an older study can be replicated. That is a legitimate limitation, not a methodological failure.
Changing the Measure Can Be More Consequential Than Changing the Setting
Not all unavoidable changes deserve equal concern.
Suppose you conduct the replication in a different classroom but use the same intervention and validated outcome measure. The setting difference may be relatively minor for some claims.
Now suppose the original outcome instrument is proprietary or unavailable, so you substitute another measure. That change may be more consequential because the dependent variable is part of the empirical claim itself.
If both instruments measure the same construct validly, the alternative may provide a useful conceptual test. But if their content, scale, sensitivity, or construct coverage differs substantially, a discrepant result becomes harder to interpret.
This is where your study may move from a close replication toward a more conceptually varied test. The distinction between direct and conceptual replication can help clarify what evidence the adapted design is expected to provide.
Changing the Analysis Can Also Change What You Are Replicating
Researchers sometimes focus heavily on reproducing materials while changing the statistical analysis without much discussion.
That can be a mistake.
If the original study tested a particular contrast or estimated a particular effect, a substantially different analytical model may estimate something else. Adding covariates, changing exclusion rules, transforming outcomes differently, replacing a frequentist test with another inferential framework, or redefining the primary endpoint can all affect comparability.
This does not mean that a flawed original analysis must always be copied. Sometimes reproducing the original analysis and adding a better-justified alternative is more informative.
Original-analysis replication Where defensible, reproduce the analysis relevant to the published claim so readers can compare the results directly.
Improved or alternative analysis Report a prespecified alternative when methodological or statistical considerations justify it.
Interpret separately Explain what each analysis estimates and whether conclusions differ.
This approach prevents a methodological improvement from quietly changing the replication target.
Some Changes Improve the Study Without Necessarily Changing the Claim
Not every departure from the original protocol should be treated as a threat.
A replication may use a larger sample, clearer documentation, preregistration, more transparent exclusion rules, improved randomization procedures, open analytical code, or stronger data-quality checks while retaining the substantive test of the original claim.
These changes can improve the credibility or precision of the new evidence without necessarily turning the study into a different research question.
However, methodological improvements should still be documented. If the new result differs from the original, readers need to know whether changes in implementation or analysis offer plausible explanations.
When Theory Is Weak, Staying Close to the Original Method Becomes More Important
There is a reason researchers often prefer close replication.
When theory and measurement are mature, researchers may be able to specify confidently which aspects of the original procedure are essential and which can vary without changing the test. When theoretical understanding is weak, that confidence is lower.
Nosek and Errington argue that methodological similarity can serve as an interim solution in such situations. By staying close to the original procedure, researchers reduce uncertainty about which methodological differences could explain discrepant outcomes.
This creates a practical rule:
The less confidently you understand which features of the original study matter, the more cautious you should be about changing them simultaneously.
If several changes are unavoidable, document them individually rather than treating the adapted study as methodologically equivalent by default.
Changing Too Much Can Turn the Study Into an Extension or a Different Test
Imagine that you cannot obtain the original intervention, use another outcome measure, recruit a theoretically different population, add a new moderator, and replace the original analytical model.
Each decision may have a defensible rationale. Collectively, however, the resulting study may no longer provide a clear retest of the original claim.
Perhaps it is a conceptual replication. Perhaps it tests generalizability. Perhaps it has become a replication-extension study. Perhaps it is simply a new study inspired by the earlier one.
The label matters less than describing the evidential relationship accurately.
If adaptations introduce additional substantive questions, consider whether you have moved into replication-extension research. If the original claim is no longer the main inferential target, calling the entire project a replication may be misleading.
Use the Two-Outcome Test Before You Call the Adapted Study a Replication
A particularly useful diagnostic comes from Nosek and Errington's claim-based definition.
Before collecting data, imagine both broad outcomes.
| Possible Outcome |
Question to Ask |
| Result is consistent with the original finding |
Would this genuinely increase my confidence in the original claim? |
| Result is inconsistent with the original finding |
Would this genuinely decrease my confidence in the original claim? |
If the answer is yes in both cases, the adapted study has a strong replication logic.
If a consistent result would be described as confirmation but an inconsistent result would immediately be dismissed because the population, measure, procedure, or setting was too different, the design provides an asymmetrical test. It may still be useful as a generalizability study, but its status as a replication is less convincing.
This test is valuable precisely because it must be applied before you know the result.
Document Differences Before Data Collection Whenever Possible
Methodological differences become much easier to interpret when they are identified before results are known.
Create a comparison between the original and proposed methods during study planning. List the population, recruitment criteria, materials, manipulation or intervention, outcome measures, setting, timing, exclusions, data-processing rules, and primary analysis.
For each difference, classify it as:
- unavoidable;
- required for ethical, legal, linguistic, technological, or practical reasons;
- an intentional methodological improvement; or
- a deliberate change intended to test generalizability or another substantive question.
Then state what you expect the difference to do. If you believe it should be irrelevant to the claim, say why. If it could affect the result, acknowledge that before collecting the data.
Preregistration can be useful for documenting these decisions and the planned interpretation, although preregistration itself does not guarantee that the replication design is appropriate.
Do Not Hide Adaptations in the Methods Section
Readers should not need to place two articles side by side and conduct forensic textual analysis to discover how your replication differed from the original.
Report consequential similarities and differences explicitly.
A comparison table can be particularly useful:
| Study Feature |
Original Study |
Replication |
Reason for Difference |
| Population |
Original eligibility and context |
New eligibility and context |
Access, generalizability, or another stated reason |
| Materials |
Original materials |
Same or adapted materials |
Availability, translation, technology, or planned variation |
| Outcome |
Original measure |
Same or alternative measure |
Availability, validity, or planned conceptual test |
| Analysis |
Published analysis |
Replication and any alternative analysis |
Comparability or methodological justification |
The actual table should contain the concrete details of your study, not generic labels. Its purpose is to make methodological departures auditable.
Do Not Decide After Seeing the Results That a Difference Suddenly Matters
Suppose your replication produces a result consistent with the original finding. You conclude that the phenomenon generalizes despite the methodological adaptations.
Now imagine exactly the same design produces an inconsistent result. Suddenly you argue that the changed instrument, population, or software makes the studies incomparable.
That is an evidential asymmetry.
Some post hoc interpretation is unavoidable because unexpected findings can reveal possibilities researchers had not anticipated. But whenever possible, identify consequential differences and their predicted implications before observing the outcome.
This is particularly important because an inconsistent replication can still provide valuable information. The question is not whether the result matches the original study but what the discrepancy permits you to learn, which becomes central when considering why a replication that does not reproduce the original finding may still be useful.