Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

UniProt: The Protein Sequence and Functional Annotation Knowledgebase

UniProt is the primary knowledgebase for protein sequence and functional annotation data, curated by a consortium of EMBL-EBI, SIB, and PIR. This guide covers UniProtKB’s Swiss-Prot/TrEMBL split, governance, licensing, and how it differs from the Protein Data Bank.

UniProt (the Universal Protein Resource) is the primary public knowledgebase for protein sequence and functional annotation data. Where a resource like the Protein Data Bank (PDB) archives the experimentally determined 3D atomic coordinates of a molecule, UniProt answers a different question: for a given protein, what is its amino-acid sequence, and what is known or predicted about what it does — its function, domains, post-translational modifications, subcellular location, disease associations, and cross-species relationships. The two resources are complementary rather than competing: most UniProt entries have no experimentally solved structure at all, and a PDB structure entry typically points back to a UniProt entry to identify which protein it represents.

What UniProt Covers, and Where It Sits Relative to PDB

For a research administrator evaluating data repositories or reviewing a data management plan (DMP) that references protein data, the distinction matters because the two resources carry different reuse obligations and different citation conventions. A study that determines a novel structure deposits coordinates in the PDB; a study that characterizes a protein’s sequence, function, or variants engages with UniProt, either as a data source (looking up existing annotation) or, less commonly for most individual researchers, as a target for structured sequence submission via the underlying nucleotide/protein sequence databases that feed it. Many UniProt entries also link out to computationally predicted structural models, including entries in the AlphaFold Protein Structure Database, for proteins that lack an experimentally solved structure — another point of convergence with, rather than substitution for, the PDB.

Governance: A Consortium, Not a Single Institution

UniProt is produced and maintained by the UniProt Consortium, a partnership of three institutions: the European Bioinformatics Institute (EMBL-EBI, UK), the Swiss Institute of Bioinformatics (SIB), and the Protein Information Resource (PIR, based at Georgetown University and the University of Delaware in the US). This tri-institutional structure mirrors the pattern seen elsewhere in life-sciences data infrastructure — the Worldwide Protein Data Bank (wwPDB) is likewise a federation of regional data centers rather than a single national archive. UniProt’s funding is similarly distributed: it draws support from the US National Institutes of Health (principally the National Human Genome Research Institute, alongside several other NIH institutes), the Swiss Federal Government, and the European Molecular Biology Laboratory (EMBL), rather than any single national science agency.

The Structure of UniProt: Four Component Databases

“UniProt” is not one database but a family of four, each serving a different purpose:

  • UniProtKB (UniProt Knowledgebase) — the core annotated protein-sequence database, and the component most researchers mean when they say “UniProt.” It is described below.
  • UniRef (UniProt Reference Clusters) — clusters sequences at set levels of sequence identity (commonly 100%, 90%, and 50%), reducing redundancy so that similarity searches and large-scale sequence analysis run faster without repeatedly re-processing near-identical sequences.
  • UniParc (UniProt Archive) — a comprehensive, non-redundant archive of essentially all publicly known protein sequences, including historical and obsolete sequences, tracked across their source databases. It functions as a sequence-level ledger rather than an annotated resource.
  • Proteomes — provides complete, non-redundant sets of protein sequences for fully sequenced genomes, organized at the organism level.

UniProtKB: Swiss-Prot (Curated) vs. TrEMBL (Automated)

UniProtKB itself is divided into two sections that differ fundamentally in how their annotation is produced:

  • UniProtKB/Swiss-Prot is the manually reviewed section. Each entry is curated by a scientist: literature is read, experimental evidence is weighed, and function, domain structure, catalytic activity, post-translational modifications, and other attributes are assigned and cross-checked against the published record. This is a comparatively small, high-confidence subset of the total sequence space, and it is the section most suited to citing as an authoritative statement about a specific protein’s known function.
  • UniProtKB/TrEMBL is the automatically annotated, unreviewed section. Sequences enter TrEMBL from translated coding sequences submitted to the underlying nucleotide sequence databases (the INSDC collaboration — GenBank, ENA, and DDBJ) and are annotated by automated pipelines and classification rules rather than by a human curator. TrEMBL is far larger in volume than Swiss-Prot and is continually updated as new sequences are submitted; its annotation quality is generally lower confidence and should be treated as computationally predicted rather than expert-verified.

This split is the single most important thing to understand about UniProt for reuse purposes: a claim sourced from Swiss-Prot carries a different evidentiary weight than the same-looking claim sourced from TrEMBL, and UniProt’s own entry pages and evidence-code system (which tags each annotation with how it was derived — experimental, curator-inferred, or automatic) make that distinction explicit and traceable.

How Annotation and Cross-Referencing Work

UniProtKB entries are extensively cross-referenced to other bioinformatics resources: structural data (PDB), gene and genome context (Ensembl, RefSeq), functional classification (Gene Ontology, InterPro, Pfam), pathway and interaction data, and disease/variant databases, among many others. This cross-referencing is what makes UniProt function as a hub rather than a standalone dataset — a single accession number becomes an entry point into a much wider web of biological data. Swiss-Prot’s manually assigned keywords and Gene Ontology terms are also used to train and validate the automatic annotation pipelines applied to TrEMBL, so the two sections are not fully independent; curated knowledge in Swiss-Prot indirectly improves the quality of automated annotation across the rest of the database.

Access, Formats, and Licensing

UniProt is freely accessible at uniprot.org, with data downloadable in multiple formats (FASTA, XML, GFF, RDF, and tab-separated text among them) and queryable programmatically through a REST API, in addition to an ID-mapping tool for translating identifiers between UniProt and other major databases. UniProt data is made available for reuse under a Creative Commons Attribution license; because license terms and specific version numbers have been revised over the resource’s history, confirm the exact current terms on the official UniProt license page before citing a specific license version in a compliance document. Each entry carries a stable accession number (for example, a code such as P12345) that functions as a persistent identifier for that protein record, allowing it to be cited and cross-referenced reliably even as annotation content is updated over time.

UniProt’s Role in Research Data Management and Compliance

For research administrators, UniProt matters less as a place researchers deposit data (most researchers are consumers of UniProt, not depositors to it directly) and more as reference infrastructure that other compliance obligations depend on:

  • DMP references. A data management plan for a proteomics, structural biology, or molecular biology project may reasonably cite UniProt as the source of reference sequence and functional annotation data, distinct from the primary experimental repository (e.g., PDB for structures, or a proteomics data repository for mass-spectrometry results) where the project’s own outputs will be deposited.
  • FAIR alignment. UniProt entries are findable via stable accession numbers, accessible via open APIs and bulk download, interoperable through extensive cross-referencing to other standard vocabularies and databases, and reusable under an open license — making UniProt a frequently cited example of infrastructure that operationalizes the FAIR Data Principles in practice.
  • Reproducibility and version awareness. Because Swiss-Prot annotation is periodically revised as new evidence emerges, and TrEMBL content changes continually, a study or grant deliverable that relies on specific UniProt annotation for a claim should record the accession number and, where reproducibility is critical, the release/version date used, rather than citing UniProt generically.

UniProt vs. the Protein Data Bank: Quick Comparison

  UniProt Protein Data Bank (PDB)
What it records Protein sequence and functional annotation Experimentally determined 3D atomic coordinates
Governance UniProt Consortium: EMBL-EBI, SIB, PIR wwPDB: RCSB PDB, PDBe, PDBj, BMRB, EMDB
Core content split UniProtKB/Swiss-Prot (curated) vs. TrEMBL (automated) Single archive of deposited experimental structures
Typical researcher relationship Mostly a data consumer / reference lookup Depositor when a structure is solved; consumer otherwise
Structural coverage Most entries have no solved structure; some link to predicted models (e.g., AlphaFold DB) Structure is the entire content of every entry

See the full Protein Data Bank (PDB) guide for PDB’s deposition mandates, format standards, and wwPDB governance in detail.

Frequently Asked Questions

Is UniProt the same thing as the Protein Data Bank?

No. UniProt covers protein sequence and functional annotation; the PDB covers experimentally determined 3D structures. A protein can have a rich UniProt entry with no PDB structure at all, and a PDB structure entry cross-references back to a UniProt entry to identify the protein it represents.

What is the difference between UniProtKB/Swiss-Prot and TrEMBL?

Swiss-Prot is manually reviewed by curators against the published literature and experimental evidence; TrEMBL is automatically annotated by computational pipelines and is not individually reviewed. Swiss-Prot is smaller and higher-confidence; TrEMBL is far larger and updated continuously.

Who governs and funds UniProt?

UniProt is produced by the UniProt Consortium: EMBL-EBI (UK), the Swiss Institute of Bioinformatics (SIB), and the Protein Information Resource (PIR, US). Funding is drawn from multiple US National Institutes of Health institutes (principally NHGRI), the Swiss Federal Government, and EMBL.

Is UniProt free to use?

Yes. UniProt is freely accessible online, with bulk downloads, a REST API, and an ID-mapping tool, and its data is released for reuse under a Creative Commons Attribution license (confirm the current license version on the official UniProt site before citing it in a compliance document).

How should I cite a UniProt entry?

Cite the specific accession number and, where reproducibility matters, the release date or version of the entry used, rather than citing UniProt generically — annotation for a given entry can change as curators incorporate new evidence.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →