Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

What Is an Institutional Repository? A Researcher’s Guide

An institutional repository is the online archive an organization runs for its own researchers’ outputs, and the core infrastructure behind green open access. This guide explains how IRs work, how deposit and self-archiving function, and how repositories differ from CRIS, subject repositories, and data repositories.

An institutional repository (IR) is an online archive, operated by a university, hospital, research institute, or other research-performing organization, that collects, preserves, and provides open access to the research outputs produced by that institution’s own faculty, staff, and students. Common contents include peer-reviewed journal articles (typically the author’s accepted manuscript rather than the publisher’s typeset version), theses and dissertations, preprints, technical reports, working papers, conference papers, and sometimes datasets or grey literature. Institutional repositories are the primary infrastructure that makes green open access possible: rather than paying a journal an article processing charge for immediate open access (gold OA), an author deposits, or “self-archives,” a copy of their work in their institution’s repository, where it becomes freely available to anyone, often after a publisher-imposed embargo period.

This guide explains what an institutional repository is, how it differs from adjacent infrastructure (subject repositories, CRIS platforms, generalist data repositories), how the underlying discovery mechanics work, and what a researcher or administrator actually needs to know to use one.

What makes a repository “institutional”

The defining feature of an institutional repository is scope by affiliation, not by subject or content type. An IR collects the output of one organization’s research community, so a single university’s repository will typically contain material spanning every discipline the institution researches — physics alongside history alongside nursing — unified only by the fact that the authors are affiliated with that institution. This is the opposite organizing principle from a subject repository, which collects material from any author, anywhere, as long as it falls within a defined discipline (arXiv for physics/math/CS, PubMed Central for biomedicine, SSRN for the social sciences).

Most institutional repositories are run by the university library, sometimes in partnership with the research office, IT services, or a research information management function. They are typically OAI-PMH compliant, meaning they expose their metadata through the Open Archives Initiative Protocol for Metadata Harvesting, a standardized protocol that lets external aggregators and search services automatically collect (harvest) records from many repositories at once without needing a custom integration for each one. This harvesting layer is what allows a single search — on Google Scholar, OAI-PMH-based aggregators like CORE or BASE (Bielefeld Academic Search Engine), or a repository directory like OpenDOAR — to surface content deposited across thousands of independently-run institutional repositories worldwide.

What an institutional repository actually does

  • Preservation — IRs are typically operated as trusted, long-term digital preservation infrastructure (fixity checking, format migration planning, backup/redundancy), distinct from a departmental file share or a personal website that may disappear when a researcher changes institutions.
  • Discoverability — each deposited item gets a persistent, citable landing page with structured metadata (title, author, date, abstract, subject terms), often with a persistent identifier such as a handle or DOI, and is exposed to harvesters via OAI-PMH so it surfaces in aggregator search results well beyond the institution’s own site.
  • Open access mechanism — by hosting a freely-downloadable copy (again, usually the accepted manuscript rather than the publisher PDF, per the publisher’s self-archiving policy), the IR is the practical vehicle for green open access compliance with funder and institutional OA mandates.
  • Institutional reporting — because deposits are tied to affiliated authors, an IR can double as a record of institutional research output for annual reports, REF/RAE-style national research assessment exercises, or internal metrics, though this function is increasingly handled by a dedicated current research information system (CRIS) instead.

How deposit and self-archiving actually work

Depositing into an institutional repository is the mechanical core of self-archiving. In practice this means:

  1. An author (or, at many institutions, a library-mediated deposit service acting on the author’s behalf) checks the publisher’s self-archiving policy to determine which version of the manuscript may be deposited — the original submitted preprint, the peer-reviewed accepted manuscript (AAM), or, less commonly, the final published version of record — and whether an embargo period applies before the file can be made public.
  2. The author uploads the file along with descriptive metadata (title, abstract, author list, funding acknowledgment, subject classification) into the IR’s deposit interface.
  3. The repository applies any required embargo, after which the item becomes openly accessible; metadata is exposed immediately for indexing even if the file itself is embargoed.
  4. The item is harvested by external discovery services and search engines, extending its reach well beyond the institution’s own repository search box.

For the complete step-by-step mechanics of this process — including how to read a publisher’s self-archiving terms, what “accepted manuscript” actually means in practice, and how embargoes and funder mandates interact — see CASRAI’s dedicated guide, Self-Archiving Research Papers for Green Open Access: A Practical How-To. The institutional repository is the piece of infrastructure that guide’s workflow deposits into; this page focuses on what that infrastructure is and how it fits into the broader research-information landscape.

Institutional repository vs. related infrastructure

The term “institutional repository” gets used loosely, but it describes a specific role that is worth distinguishing from adjacent systems a researcher or administrator will also encounter:

  • Subject (disciplinary) repositories — collect by field, not affiliation, and accept deposits from any author worldwide. See CASRAI’s Institutional Repository vs. Subject Repository comparison for a fuller breakdown.
  • Generalist data repositories (e.g. Zenodo, Figshare, Dryad) — host research data and other non-text outputs for any depositor, generally optimized for DOI-minting individual datasets rather than institutional affiliation-based collection. See Institutional Repository vs. Zenodo.
  • Current research information systems (CRIS) and research information management systems (RIMS) — track structured metadata about an institution’s research activity (people, projects, grants, publications, outputs) primarily for reporting and analytics, and increasingly integrate with or sit alongside the IR rather than replacing it. See CRIS vs Institutional Repository vs RIMS.
  • Preprint servers (e.g. arXiv, bioRxiv) — a specific case of subject repositories that host pre-peer-review manuscripts; many institutions encourage or require deposit in both a preprint server and the IR.

Common institutional repository software

Most institutions run their repository on one of a small number of established platforms rather than building custom infrastructure:

  • DSpace — open source (BSD 3-Clause license), originally built by MIT Libraries and Hewlett-Packard Labs and first launched in 2002; now stewarded by the nonprofit LYRASIS. The most widely deployed open-source IR platform, available self-hosted or through hosted options such as LYRASIS DSpaceDirect.
  • EPrints — open source (GPL), built and maintained by the University of Southampton; launched in 2000, one of the earliest OAI-PMH-based repository platforms.
  • Digital Commons — a commercial, hosted (SaaS) platform originally developed by bepress, acquired by Elsevier in 2017; bundles repository hosting with journal and open educational resource publishing tools, with no self-hosting option.
  • Fedora — open source (Apache 2.0), originally a Cornell University/University of Virginia project dating to 1997, now stewarded by LYRASIS. Fedora is a flexible object-store/API layer rather than a turnkey repository, and is typically paired with a front-end application such as Islandora or Hyrax to function as a full IR.

For a fuller feature-by-feature comparison, see CASRAI’s DSpace vs. EPrints vs. Digital Commons vs. Fedora guide.

Finding content across institutional repositories

Because individual IRs are scattered across thousands of institutions, several services exist specifically to make their combined contents searchable in one place: OpenDOAR maintains a curated directory of repositories of repositories worldwide; CORE (operated by the Open University and Jisc) and BASE (operated by Bielefeld University Library) both harvest metadata — and, where permitted, full text — from OA repositories via OAI-PMH and offer their own search interfaces and APIs; and general-purpose academic search tools such as Google Scholar also index repository content directly. A researcher does not need to know which specific institutional repository holds a paper to find an open-access copy of it — these aggregation layers exist precisely to remove that burden.

Why it matters for open access compliance

Many funder and institutional open-access policies — including the U.S. federal public access requirements now applying government-wide, cOAlition S’s Plan S, and most major university OA policies — explicitly accept green open access via repository deposit as a valid compliance route, either as the primary mechanism or as an alternative to paying for gold OA. This makes the institutional repository a piece of compliance infrastructure, not just a convenience: for a researcher trying to meet a funder mandate without an APC budget, depositing the accepted manuscript in their institution’s repository, within the funder’s required timeframe, is frequently the most direct path to compliance.

Frequently asked questions

Is depositing in an institutional repository the same as publishing open access?

Not exactly. Depositing in an IR (self-archiving) is the mechanism behind green open access — a free copy exists and is discoverable alongside the version of record, which typically remains behind the publisher’s paywall. This is distinct from gold open access, where the publisher itself makes the official published version freely available, usually funded by an article processing charge. Both routes satisfy most funder OA mandates, but they work differently.

Which version of my manuscript can I put in my institution’s repository?

It depends entirely on the publisher’s self-archiving policy for that specific journal, which typically specifies whether the preprint, the accepted manuscript, or (rarely) the published version may be deposited, and whether an embargo period applies. Publisher policies can be looked up directly (see CASRAI’s guide to self-archiving for green open access for how to check this).

Do all universities have an institutional repository?

Most research-active universities and many research institutes, hospitals, and government labs operate one, though scale and staffing vary widely — from a small library-run DSpace instance to a large, dedicated repository team. Directories such as OpenDOAR can be used to check whether a specific institution operates one.

Can an institutional repository host research data, not just papers?

Some IRs do accept datasets, especially smaller or supplementary datasets tied to a published paper, but institutions increasingly direct larger or standalone datasets to a dedicated data repository (either a generalist platform like Zenodo/Figshare or a discipline-specific one) that is purpose-built for data curation, versioning, and DOI minting. Check with your institution’s library or research data management office for local practice.

Is an institutional repository the same thing as a CRIS?

No. A CRIS (current research information system) tracks structured metadata about an institution’s research activity broadly — people, grants, projects, outputs — primarily for internal reporting and analytics, and may or may not host full-text files itself. Many institutions run both, with the CRIS feeding publication metadata into the IR, or the two systems integrated so a single deposit populates both. See CASRAI’s CRIS vs Institutional Repository vs RIMS comparison for the full distinction.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →