HEPData (the Durham High Energy Physics Database) is the primary open-access repository for scattering, cross-section, and kinematic data points extracted from published particle-physics papers. It is a discipline-specific repository in the sense CASRAI’s Dictionary defines the term: a repository scoped to one research domain’s data conventions and reuse patterns, rather than a general-purpose deposit platform for any file type.
For research-data managers and library/RDM staff who support physics faculty, HEPData is worth knowing specifically because it sits outside the generalist-repository landscape (Zenodo, Dryad) that most institutional RDM guidance defaults to. A physicist depositing a data management plan or answering a funder data-sharing requirement may point to HEPData rather than an institutional or generalist repository, and that choice is usually the field-appropriate one, not a compliance gap.
What HEPData actually archives
HEPData does not archive raw detector output or full analysis code repositories. It archives the tabulated data points, plots, and associated uncertainties that underlie a published result — the numbers behind a paper’s figures and tables, in a structured, machine-readable form (typically YAML/JSON with an accompanying schema), rather than as a static PDF table. This distinction matters for research-data compliance conversations: HEPData addresses the FAIR Data Principles‘ reusability requirement for the specific numeric results a paper reports, while raw detector data and full analysis pipelines for major collaborations (ATLAS, CMS, LHCb, and others) are typically governed separately, under each collaboration’s own open-data policy.
The record for a submission holds one or more data tables, each with its own persistent identifier (see below), plus associated plots and links back to the paper it comes from.
Host, funding, and technical basis
HEPData is hosted and operated at Durham University in the UK, with funding from the UK Science and Technology Facilities Council (STFC) supporting staff, user support, and development of the open-source software underlying the platform. The current platform (hepdata.net, replacing the earlier hepdata.cedar.ac.uk site) was rebuilt on the Invenio v3 digital-library framework — the same open-source repository framework that underlies several other research-data platforms — and the rebuild is documented in a 2017 paper by Maguire, Heinrich, and Watt, “HEPData: a repository for high energy physics data” (arXiv:1704.05473). HEPData’s own holdings extend back to data curation efforts from the 1970s, predating the current web platform by decades — a long institutional continuity that is unusual among discipline repositories and relevant context if you’re assessing HEPData’s long-term preservation credibility for a data management plan or repository-selection decision.
How HEPData connects to arXiv and INSPIRE-HEP
HEPData’s submission workflow is built around the particle-physics publication pipeline rather than treated as a standalone deposit. A submission is normally initiated using the record’s arXiv identifier once the paper is posted, or the corresponding INSPIRE-HEP record number (INSPIRE assigns this automatically once a paper appears on arXiv). INSPIRE-HEP itself is run by a collaboration of CERN, DESY, Fermilab, IHEP, IN2P3, and SLAC, and functions as the field’s shared literature/citation index; it surfaces HEPData holdings alongside the paper record so a reader moving from an arXiv preprint to its INSPIRE entry can reach the underlying data table directly. For a research-data steward, this is the practical takeaway: in particle physics, “is the data linked to the paper” is largely already solved at the field-infrastructure level, in a way many other disciplines are still building toward with generic data management plans and repository-DOI cross-referencing.
Persistent identifiers and licensing
HEPData mints DataCite DOIs for its records, separately, at both the whole-record level and the individual-data-table level — so a specific table within a multi-table submission can be cited independently of the paper’s other results. This table-level granularity is worth flagging to researchers unfamiliar with generalist repositories, where a DOI more typically covers an entire deposit rather than its component parts. The default license applied to data uploaded to HEPData is CC0, consistent with the broader open-data norm across major discipline repositories.
Where HEPData fits among other domain repositories
HEPData is one instance of a pattern CASRAI’s Dictionary already documents generally: fields with mature, decades-old data-sharing norms tend to converge on a single domain repository (see also re3data for locating the equivalent repository in other fields) rather than defaulting to institutional or generalist infrastructure. Genomics has GenBank; structural biology has the PDB; particle physics has HEPData plus the collaboration-specific open-data portals (CERN Open Data Portal, for example) for raw and derived datasets beyond publication-level tables. When advising a physics researcher on data-sharing compliance, checking whether a field-specific repository like HEPData already exists — and is the community’s actual expectation — should come before defaulting to a generalist option; funders and journals that require deposit in “a recognized repository” (per FAIR-aligned policy language) generally accept HEPData as satisfying that requirement for HEP publication data.
Practical guidance for research-administration and RDM staff
- Data management plans: where a physics DMP names a repository, HEPData (for publication-level scattering/cross-section data) is typically the field-appropriate answer rather than a generic institutional repository — confirm this with the researcher rather than substituting a generalist option by default.
- Compliance and citation: because HEPData mints table-level DataCite DOIs, encourage researchers to cite the specific table DOI, not just the paper, when their data is reused — this is both more precise and easier to track for research-output reporting.
- Scope boundary: HEPData is not the venue for raw detector data, simulation samples, or full analysis code; those follow each collaboration’s own open-data policy (e.g., the CERN Open Data Portal for ATLAS/CMS/LHCb releases) and a separate compliance conversation.
- Discoverability: because HEPData submissions are anchored to INSPIRE-HEP and arXiv identifiers, most of the recordkeeping a research office would otherwise need to do manually (linking a preprint to its dataset) is already handled by the field’s own infrastructure.
Frequently asked questions
Is HEPData the same thing as the CERN Open Data Portal?
No. HEPData holds the tabulated, publication-level results (data points, cross-sections, plots) associated with a paper. The CERN Open Data Portal and similar collaboration-run portals release raw or reconstructed detector data and simulation samples, which is a different scope, scale, and governance process.
Who can submit data to HEPData?
Submission is normally tied to a specific publication and coordinated by an author or collaboration representative associated with that paper’s INSPIRE/arXiv record, not an open, unaffiliated upload process.
Does depositing in HEPData satisfy a funder’s open-data mandate?
For publication-level HEP data, HEPData is widely treated by the field as a recognized, appropriate repository, but researchers should still confirm against their specific funder’s or journal’s data-availability policy language, since requirements vary by funder and are outside HEPData’s own control.
Is there a cost to deposit in or access HEPData?
HEPData is funded through UK STFC support to Durham University and is free to use for both deposit and access; check hepdata.net directly for any current operational specifics, since funding and operational details can change.
For the broader landscape of where domain repositories fit in a research-data strategy, see CASRAI’s research data management pillar, and for repository selection generally, CoreTrustSeal certification for research data repositories.







