03 · What You Need to Know
Duplicate Publication Is a Study-Level Problem, Not Just a Duplicate-Record Problem
First Distinguish Duplicate Records From Duplicate Publications
The word “duplicate” describes two different problems in literature reviewing.
The first occurs when the exact same article is retrieved from several databases. A paper indexed in PubMed, Scopus, and Web of Science may appear three times after search results are merged. These are duplicate bibliographic records of one report.
The second problem occurs when substantially overlapping research appears in more than one publication. The publications themselves may be different enough that reference-management software will not recognize them as duplicates.
Duplicate record
Two or more search records refer to the same publication. These are normally identified and removed during search-result deduplication.
Duplicate or overlapping publication
Two or more publications contain substantially overlapping evidence from the same underlying study, sample, or dataset.
Cochrane explicitly separates these tasks. Its study-selection workflow begins by removing duplicate records of the same report, but reviewers must later identify multiple reports that may originate from the same study.
Duplicate Publication Does Not Always Mean an Identical Article
The simplest form would be substantially the same article appearing more than once. Real duplication can be much harder to recognize.
Research examining duplicate publications used in systematic reviews has identified several patterns. Publications may contain:
- the same participants and the same outcomes;
- the same participants but different outcomes;
- a larger sample with overlapping participants and the same outcome;
- a smaller subset of the original sample;
- different reported outcomes or time points from substantially the same underlying research; or
- other forms of substantial overlap that are not obvious from the citations alone.
In a JAMA analysis of duplicate publications identified through systematic reviews, von Elm and colleagues found several distinct duplication patterns rather than one simple form. Authorship could differ substantially between related publications, which made author names alone an unreliable way of identifying overlap.
This is why duplicate-publication detection can require what Cochrane rather charmingly calls “detective work.”
Not Every Secondary Paper Is Improper Duplicate Publication
An important distinction is needed here.
A single study can legitimately generate several publications. A randomized trial might produce a protocol, primary-results article, secondary-outcome paper, economic analysis, qualitative process evaluation, and long-term follow-up. A longitudinal cohort may support many research questions over several years.
Those reports should be recognized as related, but the mere existence of several papers does not establish unethical or redundant publication.
Legitimate multiple reporting
Different reports make distinct contributions from the same research and appropriately identify or can be linked to the underlying study.
Problematic duplicate or redundant publication
Substantially overlapping material is republished in a way that may misleadingly appear to represent separate evidence, particularly when the relationship to previous publication is not made clear.
For the literature reviewer, the immediate methodological problem is similar in either case: related reports must be recognized so that the underlying evidence is handled correctly. The separate question of whether the authors committed a publication-ethics violation requires more information and should not be inferred merely because two papers share data.
Why Does Duplicate Publication Distort a Literature Review?
Imagine that five independent studies investigate an intervention. Four find little difference between groups. One reports a large positive effect.
Now imagine that the positive study generates three publications containing substantially overlapping results, and the reviewer mistakenly treats all three as independent studies.
The evidence base now appears to contain seven studies:
- four with little effect; and
- three apparently supporting a large effect.
But there were never seven independent studies. There were five.
The positive study has effectively acquired extra votes merely because its findings appeared in several publications.
Watch Out
A meta-analysis assumes that the data entering it satisfy the statistical structure required by the analysis. Counting the same participants or substantially overlapping data as independent observations can give one study inappropriate weight and distort the pooled result.
Duplicate Publication Can Bias the Direction of the Evidence
The problem would be less concerning if every study were equally likely to generate multiple publications. Evidence suggests that this is not necessarily the case.
Cochrane notes that studies with statistically significant findings or larger treatment effects may be more likely to produce multiple publications. This phenomenon is sometimes called multiple or duplicate publication bias.
That creates two related problems. Repeatedly published studies may be easier to find because several reports point to the same underlying research. If those reports are then mistaken for separate studies, their findings may also receive excessive weight in the synthesis.
Classic empirical work demonstrated that this can have substantive consequences. In one well-known investigation of trials of the antiemetic ondansetron, duplicate publication led to repeated inclusion of participant data and produced an overly favorable estimate of treatment efficacy.
The lesson is not that every multiply published study has positive findings. It is that publication multiplicity may be systematically related to results, which makes failure to identify duplicates a potential source of bias rather than merely a counting error.
How Can You Recognize Potential Duplicate Publications?
No single feature reliably proves duplication. Instead, compare several characteristics across suspiciously similar reports.
Cochrane recommends examining:
- trial registration or other study identification numbers;
- author names;
- study locations and institutions;
- specific intervention details;
- numbers of participants;
- baseline characteristics;
- study dates; and
- study duration.
Additional clues may include the same recruitment sites, identical eligibility criteria, matching demographic profiles, the same funding source, unusual intervention details, matching tables, or references to the same parent project.
| Clue |
Why it helps |
Why it is not conclusive alone |
| Same trial registration number |
Strongly links reports to one registered study |
Some reports may omit the identifier |
| Overlapping authors |
Research teams often publish several reports from one study |
The same team may conduct several independent studies |
| Same institution and setting |
May indicate a common participant pool |
Several studies can occur at the same institution |
| Similar sample size and baseline data |
Can reveal the same or overlapping participants |
Sample sizes may differ because of attrition or subgroup analyses |
| Matching recruitment dates |
Helps identify a common study period |
Parallel studies may occasionally recruit during the same period |
| Identical intervention details |
Distinctive procedures can link reports |
Standardized interventions may be reused across separate studies |
The safest conclusion usually comes from several clues pointing in the same direction.
Different Authors Do Not Prove That the Studies Are Different
This is an especially unreliable shortcut.
One report may list the principal investigator first. Another may be led by a doctoral researcher analyzing secondary outcomes. A long-term follow-up may involve a changed research team. Multicenter projects can generate papers with substantially different subsets of authors.
In the analysis by von Elm and colleagues, duplicate reports sometimes had partly or completely different authorship from their corresponding main articles.
So when participant characteristics, recruitment periods, sites, interventions, or identifiers match closely, do not dismiss possible overlap merely because the author lists differ.
Different Sample Sizes Do Not Prove Independence Either
Suppose one article reports 480 participants and another reports 417. That might look like evidence of separate studies.
It could instead reflect attrition, missing data, a subgroup analysis, an interim report, or a later follow-up.
Duplicate-publication research has documented both increasing and decreasing sample patterns among related reports. Cochrane likewise warns that participant numbers can differ across publications from the same study.
Look at recruitment dates, baseline characteristics, study sites, and participant flow before deciding that two samples are independent.
Different Outcomes Can Still Come From the Same Participants
Two publications may appear unrelated because they report completely different outcomes.
One article might examine academic performance while another examines student anxiety. If both analyses use the same participants from the same trial, they remain reports from the same underlying study.
This is not necessarily problematic. Both outcomes may legitimately contribute to a review if the review examines both.
The danger arises when the papers are counted as independent studies or when overlapping results are combined in a way that assumes independent samples.
The broader procedure for handling multiple papers from the same study or dataset therefore requires study-level organization rather than simply choosing one PDF and deleting the rest.
Duplicate Publication Is Not the Same as Salami Slicing
The terms overlap in discussions of publication ethics but should not be treated as perfect synonyms.
“Salami slicing” usually refers to dividing one research project or dataset into multiple papers, particularly when the division creates publications too narrow or overlapping to be scientifically justified. Duplicate or redundant publication more directly concerns substantial overlap with material already published.
From the perspective of a reviewer, both practices can create similar detection problems because several papers may derive from one underlying dataset.
But you do not need to determine whether authors engaged in misconduct before managing the evidence correctly. Your immediate responsibility is to identify the relationship among reports and prevent non-independent evidence from being counted as independent.
What Should You Do When You Confirm Duplication?
Do not simply delete every secondary publication.
Cochrane requires multiple reports of the same study to be collated so that the study, rather than each report, remains the unit of interest. Secondary reports may contain valuable information about study design, conduct, outcomes, or follow-up and should not automatically be discarded.
Identify the reports Determine which publications appear to originate from the same study, sample, or dataset.
Confirm the relationship Compare study identifiers, authors, sites, participant characteristics, interventions, and study dates.
Create one study-level record Link the reports under the same underlying study rather than entering them as independent studies.
Use complementary information Extract relevant methodological and outcome information from the reports that provide it.
Prevent double-counting Ensure that the same participants or outcome data do not enter a synthesis more than once as though they were independent evidence.
If the relationship remains uncertain and the distinction could affect the review, Cochrane recommends contacting report authors or principal investigators where necessary.
Which Publication Should Be Treated as the Main Report?
When several reports describe one study, reviewers often designate a primary report for organizational purposes.
Cochrane requires reviewers to choose and justify which report is used as the main source of study results, particularly when reports conflict. This does not mean the primary report becomes the only usable source.
The main report might be the publication that:
- provides the most complete results;
- reports the prespecified primary outcome;
- contains the most comprehensive participant information;
- corresponds most closely to the review's relevant time point; or
- provides the clearest account of the main study.
Other reports remain linked and can supplement it.
What If the Duplicate Reports Give Different Results?
This deserves investigation rather than an arbitrary choice.
Apparent discrepancies may result from:
- different follow-up periods;
- different participant subsets;
- adjusted versus unadjusted analyses;
- different outcome definitions;
- updated datasets;
- corrections;
- different missing-data procedures; or
- selective reporting across publications.
Check protocols, registrations, supplementary materials, corrections, and other linked reports. If the discrepancy cannot be resolved from available information, contacting the study investigators may be appropriate.
Most importantly, do not choose whichever report produces the largest, smallest, or most convenient result. The decision rule should follow the review's methods rather than the direction of the findings.
Duplicate Publication Can Affect Narrative Reviews Too
The statistical consequences are most obvious in meta-analysis, but narrative synthesis is not immune.
Suppose five papers appear to report that students respond positively to an intervention. If three of those papers come from the same underlying cohort, the literature may appear more consistently supportive than it actually is.
A narrative reviewer might write:
“Five studies reported positive experiences.”
But if those five papers represent only three independent studies, the statement overstates the amount of corroborating evidence.
Counting independent evidence matters even when no pooled effect size is calculated.
Duplicate Publications Can Also Distort Perceptions of Replication
Replication is especially important because repeated findings from independent samples provide different evidence from repeated analyses of the same participants.
If three publications from one dataset all reach similar conclusions, that may show internal consistency across analyses. It does not demonstrate that three independent research teams reproduced the finding in three independent populations.
Repeated publication or analysis
Several reports or analyses derive wholly or partly from the same underlying participants or data.
Independent replication
A separate study collects independent evidence capable of corroborating, refining, or challenging the original finding.
A literature review should not accidentally transform the first into the second.
Study-Level Tracking Helps Prevent the Problem
The simplest practical safeguard is to track both reports and underlying studies.
Give each study a unique study ID. Then link every related publication to that ID.
For example:
Study ID: DIGITAL-FEEDBACK-07
- Report A: conference abstract;
- Report B: primary journal article;
- Report C: secondary-outcome article; and
- Report D: follow-up publication.
This structure makes it much harder to accidentally count four reports as four studies. It also helps when preparing study characteristics, extracting outcomes, documenting exclusions, and constructing the final study-selection flow.
PRISMA Distinguishes Reports From Studies
PRISMA 2020 distinguishes records, reports, and studies in study-selection reporting. This allows a review to report, for example, 60 included reports representing 45 included studies.
That distinction is useful precisely because the relationship is not always one paper to one study.
When a PRISMA flow diagram is appropriate for your review, preserving study-level and report-level identities during screening makes the final reporting considerably easier.