03 · What You Need to Know
Where Legal Open Copies Come From and How to Find Them
First, understand what you are actually looking for
When researchers say they want a "free copy" of a paper, several different things can be meant. The journal's published version might itself be openly accessible. Alternatively, another legitimate version of the work may have been deposited somewhere else.
This distinction matters because scholarly articles can exist in multiple versions. Depending on the publication history and the rights governing the work, you might encounter a preprint, an author-accepted manuscript, or the final published version.
Open publisher version
The article is available without a subscription from the publisher or journal itself.
Repository version
A version of the work has been deposited in an institutional, subject, or other scholarly repository and made publicly accessible.
The repository copy is not necessarily inferior simply because it lacks the publisher's typesetting. The more important questions are what version it represents and whether any differences from the final publication matter for your purpose.
Start with Google Scholar
Google Scholar is one of the simplest places to begin because it attempts to identify multiple versions and access routes for scholarly works. Its official help documentation recommends several strategies for locating full text: use a library link when one appears beside the result, follow a PDF link displayed to the right, and inspect All versions for alternative sources.
If an ordinary keyword search produces too many results, search the complete article title in quotation marks. Google Scholar specifically recommends quotation marks for title searching.
Search the exact title Put the complete article title in quotation marks when necessary to reduce unrelated results.
Look beside the result Check for a PDF or HTML link that leads to an accessible version.
Open All versions Inspect alternative locations associated with the same scholarly work.
Verify what you found Confirm that the authors, title, publication details, and manuscript itself correspond to the paper you intended to retrieve.
Google Scholar groups multiple versions of scholarly works when it can identify them as versions of the same work. These may include preprints and other manifestations in addition to the publisher version.
Do not assume, however, that the first search result will expose every accessible copy. Discovery systems have different coverage, and repositories are not always indexed perfectly.
Use the DOI as a precise identifier
If the paper has a Digital Object Identifier, or DOI, keep it handy. A DOI identifies a scholarly object much more precisely than a few words from its title, particularly when the title is short or generic.
You can copy the DOI from the publisher page, database record, reference list, or article metadata and use it when searching open-access discovery services. It is also useful for checking whether a repository record corresponds to the publication you need.
Do not confuse the DOI with evidence that a particular PDF is authorized or identical to the published version. The DOI helps establish the identity of the scholarly work. You still need to inspect the source and version of the file you find.
Check Unpaywall
Unpaywall is specifically designed to locate open-access scholarly articles. According to its documentation, it harvests content from more than 50,000 journals and open-access repositories and incorporates information from sources including PubMed Central, the Directory of Open Access Journals, Crossref, and DataCite. It searches for scholarly articles using DOIs.
Its browser extension can recognize a DOI on a webpage and query Unpaywall for an open-access location. This makes it particularly convenient when you are already looking at a paywalled publisher page.
For researchers concerned about whether "open" really means legally available, the distinction is important. Unpaywall states that it harvests from legal sources, including repositories operated by universities, governments, and scholarly societies, as well as open content hosted by publishers.
That makes it fundamentally different from simply searching the wider web for a filename ending in ".pdf."
Search institutional repositories
A university repository is another important place to look, particularly when you know where one of the authors works or worked when the research was conducted.
Institutional repositories are publicly accessible scholarly collections managed by universities or other research organizations. Depending on applicable copyright, publisher, funder, and institutional policies, researchers may deposit versions of their publications there. The deposited copy is often an earlier manuscript or an author-accepted manuscript rather than the final publisher PDF.
A targeted web search can sometimes locate these records:
- search the exact paper title together with the author's university;
- search the title together with the word "repository";
- search the DOI together with the author's institution;
- visit the institution's repository directly and search by author or title.
If you are unfamiliar with these systems, understanding what an institutional repository contains and why papers appear there makes this search considerably easier.
Search subject repositories when they are relevant to your discipline
Some fields have well-established repositories or preprint servers used by their research communities. Depending on the discipline, these can contain manuscripts that precede formal journal publication, accepted manuscripts, or other openly shared research outputs.
The appropriate repository depends on the field, and repository practices differ substantially among disciplines. A physicist, an economist, a biomedical researcher, and an education researcher should not assume that the same repository ecosystem applies to all four.
This is also why a general web search should not be your only strategy. Scholarly repositories provide provenance and metadata that can make it easier to determine what a document is and where it came from.
Use an open-access aggregator such as CORE
Searching individual repositories one by one can become tedious. Repository aggregators address that problem by bringing records from many sources into a common discovery system.
CORE describes itself as a not-for-profit service that aggregates open-access research from repositories and journals. Its infrastructure harvests metadata from repositories and other providers, allowing open research distributed across many systems to become searchable through a common service.
This can be especially useful when you suspect that an open manuscript exists but do not know which institution or repository holds it.
Do not expect every open copy to look like the journal PDF
This is where researchers sometimes abandon a perfectly useful source too quickly. You search for a paper and find a plain-looking manuscript with no journal logo, no polished two-column layout, and perhaps no final page numbers. It may nevertheless be a legitimate version of the article.
Repository-based open access often involves self-archiving an earlier manuscript or the author-accepted manuscript. Publisher policies governing what authors may deposit, where they may deposit it, and whether an embargo applies vary considerably.
An author-accepted manuscript generally represents the manuscript after peer-review revisions but before the publisher's final production process. Whether that is sufficient depends on what you need from the paper. If you are extracting substantive findings, it may be entirely useful. If you need final pagination, corrected wording, supplementary files, or the definitive published record, the distinction can become more consequential.
Before deciding that a repository copy is "not the real paper," determine whether the difference between the accepted manuscript and published version matters for your use.
Verify that the open copy is actually the same work
Finding a document with a familiar title is not the end of the process. Verify its identity before using it.
Compare the open copy against the publisher record or another reliable bibliographic record. Check the title and authors first. Then compare the journal, year, volume, issue, DOI, and other available publication details.
Some differences are normal. A preprint may have a different title from the final article. Author order can occasionally change during publication. Pagination may be absent from a manuscript version. The point is not that every metadata field must look identical, but that you should have sufficient evidence that the document represents the work you intended to retrieve.
Then identify which version you have
Once identity is established, determine the version. Look for labels such as "preprint," "submitted manuscript," "accepted manuscript," "author accepted manuscript," "post-print," or "version of record." Repository metadata may provide this information even when the PDF itself does not.
| Version you may encounter |
What it generally represents |
What to check |
| Preprint or submitted manuscript |
A version from before formal peer-review revisions were completed |
Whether substantive content changed before publication |
| Author-accepted manuscript |
The manuscript after acceptance and peer-review revisions but before final publisher production |
Whether production changes, pagination, figures, supplements, or corrections matter for your purpose |
| Version of record |
The formally published version maintained by the publisher |
Whether later corrections, updates, or retractions exist |
Version awareness is part of responsible source use. "I found the paper" and "I found the final published version" are not always the same statement.
What makes an open copy a safer choice?
There is no single visual feature that proves legality. Instead, provenance matters. Copies hosted openly by the publisher or deposited in established institutional, government, scholarly society, or disciplinary repositories provide a clearer basis for evaluating where the file came from.
Open-access discovery services can further reduce uncertainty by deliberately indexing established open sources. Unpaywall, for example, explicitly limits its discovery approach to legal sources.
Watch Out
A PDF appearing in a search engine does not, by itself, establish that the file was lawfully shared, that it is complete, or that it represents the final publication. Verify the host, bibliographic identity, and manuscript version rather than treating "free to download" as synonymous with "verified open access."
What if the only copy you find is on ResearchGate or Academia.edu?
Academic networking platforms can lead you to manuscripts uploaded by researchers, but the presence of a document on a platform does not by itself tell you what rights govern that particular upload or which manuscript version it contains.
Rather than assuming that every file on such a platform is either legitimate or illegitimate, examine the specific document and its provenance. The more useful question is what you should verify before trusting a paper uploaded to an academic networking platform.
If no legal open copy appears, change access routes
Open-access discovery does not guarantee success. Some articles simply do not have an openly available version, or an existing manuscript may not yet be publicly available because of an embargo or other restriction.
If your searches fail, return to the broader set of options for accessing a paper that remains behind a paywall. Your institutional library may subscribe to the journal, provide document delivery or interlibrary loan, or help locate an accessible copy that your own searches missed.
You can also ask the author whether they can share the paper directly, subject to whatever rights and permissions apply to the work.