03 · What You Need to Know
Duplicate References Are Usually a Workflow Problem Before They Are a Software Problem
A duplicate reference exists when your library contains more than one bibliographic record representing the same work when one authoritative record would be sufficient for your purposes.
That definition matters. Seeing the same source in several collections does not necessarily mean it has been duplicated. In Zotero, for example, one item can belong to multiple collections without creating additional library records. Zotero compares this model to adding one song to several playlists. A source appearing under both “Dissertation” and “Journal Article” may therefore still be one item rather than two.
One record in multiple collections
The same library item is associated with several projects or organizational categories. This is not necessarily duplication.
Multiple records for the same work
Independent bibliographic items represent the same underlying work. These may need to be reviewed and consolidated.
Why duplicate references appear
Most duplicates are not mysterious. They arise because researchers encounter the same literature repeatedly.
You might import a paper from Scopus during one search and encounter it again in Web of Science during another. A collaborator may add a source already present in a shared collection. You may import a bibliography containing records that overlap with your existing library. A PDF may be added independently even though its bibliographic record already exists. Long-running research programs make these situations increasingly likely because the same literature often contributes to several projects.
Duplicate records can also differ enough to escape immediate recognition. One database may abbreviate a journal title while another provides the full title. Author names may be represented differently. One record may contain a DOI while another does not. Titles can differ in capitalization, punctuation, or special characters.
This is one reason accurate metadata matter. Poor metadata do not merely produce citation problems. They can also make records harder for both humans and software to recognize as representing the same work.
Use one record across projects instead of importing the paper repeatedly
A simple preventive principle is to treat your reference library as a persistent research database rather than a separate bibliography for every project.
If a paper already exists in your library, reuse that record. Associate it with another project collection, add a relevant tag, or otherwise organize it for the new context instead of importing a second copy merely because the paper has become relevant somewhere else.
This is particularly important when you use the same reference library across multiple research projects. One well-maintained record can serve several manuscripts without requiring parallel copies.
How your software implements this varies. Zotero explicitly allows one item to belong to multiple collections without duplication. Other reference managers provide their own organizational models. Understand that model before assuming that every project needs an independent copy of every source.
Search before importing when you suspect you already have the paper
You do not need to perform an exhaustive library search before saving every article. That would turn reference management into a remarkably effective method for avoiding actual research.
But when a paper looks familiar, a quick search can prevent unnecessary duplication. Search the title, DOI, author, or another distinctive field before importing it again. This is especially useful for foundational papers, frequently cited methodological sources, reporting guidelines, and other works you encounter repeatedly.
Persistent identifiers can help here. A DOI can provide a strong indication that two journal-article records refer to the same work. Mendeley Reference Manager, for example, identifies two references with the same DOI as duplicates. Zotero's current duplicate-detection algorithm considers title, DOI, and ISBN, along with publication year and creator information under specified conditions.
Identifiers are useful but not universal. Not every legitimate source has a DOI, and bibliographic records may omit identifiers that actually exist. Understanding the role of DOIs, PMIDs, ISBNs, and other research identifiers can make both retrieval and record verification more reliable.
Do not create a second bibliographic record merely to attach another PDF
A common source of clutter is treating every PDF as though it needs a new reference record.
If the bibliographic item already exists, determine whether the new file belongs with that existing record. Depending on your reference manager, an item may support attachments without requiring another copy of the bibliographic metadata.
This distinction becomes particularly important when deciding whether to attach PDFs to your reference manager. Your file-management decision should not accidentally create multiple bibliographic identities for one source.
Run duplicate detection periodically instead of waiting for a crisis
Even careful researchers will accumulate duplicates. Reference managers therefore provide tools for identifying likely matches.
Zotero has a special Duplicate Items collection. Its current documentation states that duplicate detection considers title, DOI, and ISBN and, depending on the available information, also compares publication years and creator lists. Potential duplicates can then be reviewed and merged.
Mendeley Reference Manager provides a Duplicates smart collection. Mendeley's documentation explains that references sharing the same DOI can be identified as duplicates and grouped into duplicate sets for review.
EndNote also provides duplicate-detection functionality. Its documentation describes tools for identifying duplicate records using bibliographic fields, with the exact workflow and comparison criteria depending on the EndNote environment and settings.
The important point is that these systems identify potential duplicates. They do not eliminate the need to inspect what has been matched.
Why duplicate detection can miss genuine duplicates
Software must compare something. If the records disagree on the fields used for comparison, the same work may not be recognized automatically.
Consider two records:
| Field |
Record A |
Record B |
| Title |
Artificial Intelligence in Higher Education: A Review |
Artificial intelligence in higher education - a review |
| Author |
Garcia, Manuel B. |
Garcia, M. |
| Year |
2025 |
2025 |
| Journal |
Journal of Example Research |
J. Example Res. |
| DOI |
Present |
Missing |
A researcher can probably recognize these as possible versions of the same record. Whether software does so depends on its matching logic and the metadata available.
That is why correcting inaccurate citation metadata contributes to more than prettier bibliographies. Consistent records can also make library maintenance easier.
Why duplicate detection can also produce false positives
The opposite problem is equally important. Similar metadata do not prove that two records should be combined.
An author may publish a conference paper and later a substantially developed journal article with a similar title. A preprint may precede a version of record. A book may have several editions. A correction or erratum may resemble the original publication. Articles in a series may have nearly identical titles.
These relationships need interpretation. If the records represent distinct citable works, merging them merely because the titles look alike can destroy meaningful bibliographic distinctions.
Watch Out
Never treat a reference manager's duplicate list as an instruction to merge everything it displays. It is a set of candidate matches. Confirm that the records actually represent the same work before consolidating them.
When duplicates really are duplicates, merging is often safer than simply deleting
If two records represent the same work, you still need to decide how to consolidate them.
In Zotero, the official recommendation is explicit: resolve duplicate items by merging them rather than deleting one copy. Zotero states that merging preserves collections and tags from the merged items and is recognized by its word-processor plugins, helping existing automatically generated citations and bibliographies continue to function.
Zotero's merge interface also allows you to select a master item and choose among conflicting field values. This matters because duplicate records are often complementary rather than identical. One may have the correct DOI while another has better author information. One may contain useful tags while the other belongs to the correct collections.
Mendeley Reference Manager handles duplicates differently. Its current duplicate-management workflow groups detected records in a Duplicates smart collection and provides options for retaining the desired reference while removing unwanted duplicate records. The exact operation therefore depends on the software you use rather than on a universal “merge” command.
Before resolving duplicates, understand what your reference manager preserves and what it discards.
Compare more than the title before choosing what to keep
When two records are confirmed duplicates, compare their contents before resolving them.
Check the authors, title, publication year, journal or publisher, volume, issue, pages or article number, DOI or other identifier, and item type. Then inspect information specific to your own library: tags, collections, notes, attachments, annotations, links, and citation relationships.
Do not automatically retain the record that looks more complete. Completeness and correctness are not the same thing. A record can contain many fields and still contain the wrong data.
When uncertain, verify the bibliographic information against an authoritative source such as the publisher's record, Crossref where appropriate, PubMed for relevant biomedical literature, or another authoritative bibliographic source for the work.
Duplicate records can create citation problems
Duplicates are not merely cosmetic.
If you cite different copies of the same work in a manuscript, your reference manager may treat them as independent bibliographic items. Zotero documents, for example, that citing duplicate items separately can trigger author-name disambiguation because the citation processor interprets them as different sources.
More generally, duplicate records can leave one citation linked to a record you later delete while another citation points to the surviving copy. This is one reason Zotero recommends merging rather than deleting duplicate items. Zotero also notes that deleting and reimporting an item can leave existing citations disconnected from the active library record.
Cleaning duplicates before intensive manuscript writing is therefore easier than discovering the problem while troubleshooting an apparently inexplicable bibliography.
Bulk imports deserve special attention
Duplicate accumulation accelerates when you import large result sets from multiple databases.
Suppose you search Scopus, Web of Science, PubMed, and another database for a review. The same article may appear in several result sets. In that context, duplication is expected rather than exceptional.
For systematic reviews, scoping reviews, and other evidence syntheses, deduplication is part of a documented study-selection workflow. The objective is not simply to keep a personal reference library tidy. Records may need to be imported, deduplicated, screened, counted, and reported in accordance with the review methodology.
Do not casually delete overlapping search records before you understand how deduplication fits into the review protocol and the software used for screening. A personal library cleanup and a reproducible evidence-synthesis deduplication process are not the same task.
Shared libraries need agreed practices
Collaboration creates another route to duplication. Two researchers can independently add the same paper because neither realizes the other has already imported it.
If a team uses a shared library, establish a simple convention. Team members might search the shared library before importing known foundational sources, use consistent metadata practices, and assign responsibility for periodic duplicate review.
A complicated governance manual is rarely necessary. The objective is merely to avoid having several people repeatedly create parallel records for the same literature.
Do not repeatedly export and reimport your own library
Export and import are useful for migration and interoperability, but they should not become a routine way of synchronizing the same reference library across computers or projects.
Reference managers commonly provide dedicated synchronization mechanisms. Zotero, for example, distinguishes data syncing from file syncing and synchronizes library items, notes, links, and tags through its data-syncing system.
Repeatedly exporting a library and importing it back into another copy can create unnecessary opportunities for duplicate records, particularly when the receiving library already contains some of the same items.
If your objective is to maintain one library on several devices, use the platform's supported synchronization workflow rather than treating each device as an independent database that must continually be recombined.