03 · What You Need to Know
Why Universities Keep Repositories and What You Can Find in Them
An institutional repository is more than a folder of uploaded PDFs
The concept of an institutional repository includes both a digital collection and the services surrounding it. A widely cited formulation by Clifford Lynch describes a university institutional repository as services offered to members of the university community for managing and disseminating digital materials created by the institution and its members. The concept also involves organizational stewardship, including preservation, organization, and access.
That broader role is important. A repository is not merely somewhere a researcher happens to have uploaded a file. Institutional repositories are typically connected to a university, research organization, library, or consortium and are designed to manage scholarly or intellectual outputs systematically.
The American Library Association similarly describes an institutional repository as a digital repository for collecting, preserving, and disseminating an organization's informational or research output.
Repositories can contain much more than journal articles
Although you may encounter an institutional repository while searching for a journal article, repositories commonly hold several kinds of scholarly and institutional materials.
Depending on the repository's scope, these can include journal manuscripts, conference papers, theses and dissertations, reports, books or book chapters, datasets, teaching materials, presentations, and other outputs produced by faculty, researchers, staff, and students.
Repository collections differ considerably among institutions. One university may emphasize faculty publications and theses. Another may include datasets, conference materials, institutional publications, historical collections, or other locally produced works.
Do not assume that every university repository has the same deposit rules, content types, or access policies.
Why would a journal paper be in a university repository?
One major reason is self-archiving. An author may be able to deposit a version of a journal article in an institutional repository under the conditions applicable to that publication.
This creates an important distinction. The journal and the repository do not necessarily host the same file under the same access conditions. A publisher may restrict access to its final published version while permitting an author-accepted manuscript to become publicly available through a non-commercial institutional repository, sometimes after an embargo period.
For example, Elsevier's current hosting policy allows institutional repositories to host authors' accepted manuscripts publicly after the applicable embargo period, subject to specified conditions. Its policies distinguish that manuscript from the published journal article.
This is one reason a paper can legitimately appear in a repository even though the publisher's own copy remains behind a paywall.
A repository is not necessarily giving away the publisher PDF
This is perhaps the most important point to understand when you encounter the same paper in two places.
Scholarly articles pass through several versions during publication. A repository may contain a version the author is permitted to deposit rather than the final file distributed by the publisher.
| Version |
What you may see in a repository |
What it means for you |
| Preprint |
An author manuscript from before formal journal peer review and acceptance |
Useful for access to the research, but it may differ substantively from the eventual published article |
| Accepted manuscript |
The manuscript accepted after peer-review revisions where applicable, without the publisher's final production work |
Often contains the substantive peer-reviewed content but may differ from the version of record |
| Published version |
The final version of record, when its license or other applicable rights permit repository distribution |
Corresponds to the formally published article, but you should still check for later corrections or updates |
The exact terminology can vary among publishers and repositories. If you find an accepted manuscript, understanding how it differs from the published version will help you decide whether that version is sufficient for what you need to do.
Why would an author deposit an accepted manuscript?
The accepted manuscript occupies a useful position in the publication process. It typically contains revisions resulting from peer review and editorial assessment but precedes publisher additions such as copyediting, typesetting, final formatting, technical enhancements, and pagination.
Depending on the publishing agreement and applicable policies, authors may retain or receive rights to archive this version even when public distribution of the final publisher file is restricted.
Elsevier, for example, permits accepted manuscripts to be hosted publicly in non-commercial institutional repositories after journal-specific embargo periods and under specified conditions. That is an example of a publisher policy, not a universal rule. Other publishers, journals, licenses, institutions, and funders may impose different conditions.
This is why the correct question is not simply, "Is this paper copyrighted?" Almost all contemporary scholarly articles involve copyright. The more useful question is whether the particular version has been deposited and made accessible under the rights and conditions that apply to it.
What is an embargo?
An embargo is a period during which a deposited manuscript is not yet made openly available to the public under a particular policy.
A repository record may therefore exist before its full text becomes publicly accessible. You might see the title, authors, abstract, DOI, and other metadata but encounter a notice indicating that the attached manuscript will become available on a later date.
Embargo lengths and conditions vary. Elsevier, for example, states that its journal-specific embargo periods for accepted manuscripts vary and are commonly in the range of 12 to 24 months, with exceptions. Other publishers may use different periods or conditions.
Do not generalize one publisher's embargo rules to another. If access is restricted, inspect the repository record and the relevant publisher or journal policy rather than guessing when the manuscript should become available.
The repository record can be as useful as the PDF
When you find a paper in an institutional repository, do not immediately click the download button and ignore everything else on the page.
The repository record may contain information that helps establish exactly what you have found. Depending on the system, this can include:
- the article title and authors;
- publication year;
- journal or conference information;
- DOI or another persistent identifier;
- abstract and keywords;
- document or manuscript version;
- deposit or availability information;
- license or rights information;
- a link to the formally published article.
Those metadata help connect the repository object to the wider scholarly record. A DOI link is particularly useful because it can take you to the publisher's current record even when the full published text remains inaccessible.
A repository can preserve research as well as make it discoverable
Access is only one purpose of an institutional repository. Preservation is another.
Research libraries have long treated repositories as infrastructure for organizing and preserving institutional scholarship. Repository services can apply metadata, manage digital files, support discovery, and maintain scholarly materials beyond the lifespan of a researcher's personal website or local storage.
This distinction matters because personal web pages are fragile. Researchers change universities. Departmental websites are redesigned. File paths disappear. A repository is intended to provide institutional stewardship rather than depend entirely on one researcher's continued maintenance of a webpage.
Repositories can also increase discoverability
Institutional repositories do not exist in isolation from scholarly search systems. Repository metadata can be exposed for harvesting and indexing, allowing records to become discoverable through external services.
This helps explain why a Google or Google Scholar search may lead you to a repository that you had never heard of. You did not necessarily search that university intentionally. The repository's metadata or full text may have been indexed by a discovery service.
Open repository infrastructure also supports aggregation across institutions. Services can harvest repository metadata and make outputs distributed across many universities searchable through larger discovery systems.
For a researcher trying to find a legal open-access version of a paper, this distributed infrastructure is particularly valuable. You may discover a manuscript without knowing in advance which institution holds it.
Institutional repositories and subject repositories are not the same thing
Both can provide access to scholarly outputs, but they are organized around different principles.
Institutional repository
Organized around the scholarly and intellectual output of a particular university, research institution, organization, or consortium.
Subject repository
Organized primarily around a discipline or research community and may contain work from researchers at many institutions.
arXiv, for example, is organized around subject communities rather than the output of one university. Europe PMC similarly serves a disciplinary research ecosystem. An institutional repository, by contrast, is tied to the outputs and stewardship responsibilities of an institution or group of institutions.
Neither model is inherently better for finding a particular paper. The appropriate place to search depends on the discipline, author, publication, and deposit practices involved.
Is a paper trustworthy because it is in a university repository?
A repository provides useful provenance, but repository location should not be confused with peer review or research quality.
A university repository can contain peer-reviewed accepted manuscripts, but it may also contain preprints, theses, conference presentations, reports, datasets, teaching materials, and other outputs. The institution hosting the file does not thereby certify every item as peer reviewed, methodologically sound, or correct.
Evaluate the scholarly status and quality of the work itself. If the item is a journal manuscript, determine whether it is a preprint, accepted manuscript, or published version. Then assess the research using the same critical standards you would apply elsewhere.
Watch Out
"Hosted by a university" and "peer reviewed" are not synonymous. An institutional repository tells you something useful about the provenance and stewardship of a file, but you still need to identify the document type, publication status, and version before deciding what evidential weight to give it.
How can you tell whether the repository copy is really the paper you want?
Compare the repository record with the authoritative bibliographic record for the article. Check the title and authors, then the journal, year, volume, issue, DOI, and other details when available.
If the title differs slightly, do not immediately conclude that it is a different paper. Titles can change between preprint, accepted, and published versions. Instead, use the complete metadata and manuscript content to establish the relationship.
Follow the DOI or publisher link supplied by the repository. Even if the publisher page is paywalled, it can help confirm the article's bibliographic identity and reveal whether corrections, updates, or retraction information have subsequently been attached to the publication.
What if the repository contains only a record and no downloadable file?
Not every repository record provides immediate public full-text access. The manuscript may be under embargo, restricted for rights reasons, unavailable for deposit, or represented only by bibliographic metadata.
A metadata-only record can still help. It may confirm the paper's identity, provide a DOI, identify the author and institution, or show the manuscript's access status.
If you still cannot obtain the full text, return to the broader options for accessing a paper that remains behind a paywall. Your library may provide access or document delivery, and you may also be able to request a copy directly from the author.