Skip to main content
v2026.11,772 entries · CC-BY 4.0

Genomics England’s National Genomic Research Library: How Researchers Get Access

How researchers apply to access Genomics England’s National Genomic Research Library through the governed Research Environment, and what’s actually in it.

Ask about Genomics England’s National Genomic Research Library: How Researchers Get Access

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Genomics England’s National Genomic Research Library (NGRL) is the researcher-facing dataset behind the 100,000 Genomes Project, and it keeps growing years after that project itself finished recruiting — ongoing NHS Genomic Medicine Service (GMS) participants are still being added. This page covers what the NGRL actually contains, how a researcher applies to use it, and the access model it runs on, which is a real, working example of the “compute-to-data” trusted research environment (TRE) approach that CASRAI’s own Secure Data Enclave and Data Trusts pages describe more generally.

What is actually in the library

As of the most recent published figures, the NGRL contains more than 140,000 whole genomes drawn from NHS patients with cancer or rare conditions, plus their relatives where relevant (probands and family members for rare-disease cases; matched germline and somatic genomes for cancer cases). Alongside the raw sequencing data sit alignments and variant calls, phenotypic data, and longitudinal medical history drawn from linked hospital records, cancer treatment datasets, and mental health datasets — with proteomic, transcriptomic, and digital histopathology data being added as newer, emerging dataset types. Participants come from two sources: the original 100,000 Genomes Project cohort, and ongoing recruitment through the NHS Genomic Medicine Service, which means the library is a living, continuously expanding resource rather than a fixed, closed dataset.

This is a deliberately different thing from the 100,000 Genomes Project itself, which was a time-bound recruitment programme (2013–2018) that established the initial cohort. The NGRL is the ongoing data infrastructure that project fed into and that NHS GMS keeps feeding — the distinction matters because searches for “100,000 Genomes Project” are largely about the historical programme and NHS diagnostic pathway, while a researcher planning an analysis needs the NGRL and Research Environment specifically.

How a researcher gets access

Access is not open or self-service. A researcher applies to join the Genomics England Research Network as either an academic or an industry researcher; approval is required before any data access is granted. Governance sits with an independent Access Review Committee, which sets the criteria for who can access the library and what uses of the data are acceptable — the stated principle behind the review process is that protecting participant data, which was contributed under specific consent terms, takes priority over convenience of access. This is a materially different model from a self-service open-data repository: every approved use is reviewed against a specific research purpose, not granted as a blanket credential.

The Research Environment: analysis without data export

Once approved, a researcher does not download the data. All analysis happens inside Genomics England’s cloud-based Research Environment, a trusted research environment built on Amazon WorkSpaces virtual desktops. Individual patient-level data cannot leave the environment — a researcher can export analysis results and outputs, but not the underlying genomic or clinical records themselves. Practically, getting into the Research Environment means:

  • Completing Information Governance training as part of onboarding, before or alongside first login
  • Installing the Amazon WorkSpaces client (or using the browser-based option) and registering with the access code issued by Genomics England on approval
  • Authenticating via two-factor authentication through Okta on every session, in addition to a standard email/password login

The Research Environment provides a range of analysis tools intended to serve both coding researchers (command-line bioinformatics tools, Python, R) and non-coding researchers, though the documentation is explicit that meaningful large-scale analysis of whole-genome data generally does require programming competency. Genomics England publishes detailed technical documentation for the environment separately from its general public-facing pages, reflecting that the actual user base is working researchers rather than the general public.

Why this model matters beyond genomics

The NGRL/Research Environment pairing — a governed access-approval process feeding a locked-down compute environment that never releases raw records — is one of the UK’s most mature working examples of the TRE model that other national data infrastructure is now being built around, including the components of the UK’s planned National Data Library and the parallel Health Data Research Service. It sits alongside Health Data Research UK‘s own federated data-access work and the community-governance models CASRAI documents in its coverage of H3Africa’s Data and Biospecimen Access Committee — a structurally similar review-committee-gated access model built for a different population and dataset. For a data-management-plan author writing about how sensitive human genomic data will be shared, “access via a governed TRE, no raw-record export” is the specific, verifiable pattern the NGRL demonstrates, rather than a generic promise to “share data responsibly.”

Frequently asked questions

Is the National Genomic Research Library the same as the 100,000 Genomes Project?

No. The 100,000 Genomes Project was the time-bound programme that built the initial cohort (2013–2018); the NGRL is the ongoing dataset that project populated and that the NHS Genomic Medicine Service continues to add participants to.

Can data from the NGRL be downloaded or exported?

No. All analysis happens inside the Research Environment; researchers can export their analysis outputs and results, but not the underlying patient-level genomic or clinical records.

Who can apply for access?

Both academic and industry researchers can apply, by joining the Genomics England Research Network; every application is reviewed against the Access Review Committee’s criteria rather than granted automatically.

Sources: Genomics England, “The Research Environment” (genomicsengland.co.uk/research); Genomics England, “The National Genomic Research Library” (genomicsengland.co.uk/research/data); Genomics England Research Environment technical documentation, “Accessing the RE” (re-docs.genomicsengland.co.uk).

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.