Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

GenBank: NCBI’s Foundational DNA and Genetic Sequence Database

GenBank is NCBI’s open, annotated archive of all publicly available DNA and RNA sequences — the base nucleotide layer that many other genomics repositories build on or reference. This guide covers what GenBank is, how it differs from RefSeq, how to submit and cite a sequence, and how it relates to the more specialized repositories already covered in this cluster.

GenBank is the United States’ national archive of publicly available DNA and RNA sequences, maintained by the National Center for Biotechnology Information (NCBI), part of the National Library of Medicine at NIH. It is, in a meaningful sense, the base layer underneath most of the genomics infrastructure discussed elsewhere on this site: a raw, annotated collection of nucleotide sequences — from a single gene submitted by one lab to entire assembled genomes — organized by accession number and cross-referenced by taxonomy, rather than by a specific experiment type or assay.

That distinguishes it from the more specialized repositories already covered in CASRAI’s RDM cluster. GEO holds gene-expression experiments, dbGaP and EGA hold individual-level genotype-phenotype data under controlled access, ArrayExpress holds functional-genomics assay data, and Metabolomics Workbench, MetaboLights, and ProteomeXchange cover metabolomics and proteomics respectively. GenBank sits underneath nearly all of them: a gene-expression study deposited in GEO, a variant call deposited in dbGaP, or a genome assembly cited in a paper will very often point back to a GenBank (or RefSeq) accession number for the underlying sequence itself. This guide covers GenBank on its own terms — what it is, how it differs from NCBI’s curated RefSeq database, how researchers submit and cite sequence data, and where it fits relative to the rest of this cluster — rather than repeating ground already covered by CASRAI’s guide on genome projects and research data management, which covers GenBank’s historical role as the deposit target of the Bermuda and Fort Lauderdale data-release principles.

What GenBank Is, and How It Fits Into INSDC

NCBI describes GenBank as “the NIH genetic sequence database, an annotated collection of all publicly available DNA sequences.” It is not a US-only silo: GenBank is one of three synchronized databases forming the International Nucleotide Sequence Database Collaboration (INSDC), alongside the European Nucleotide Archive (ENA, maintained by EMBL-EBI) and the DNA Data Bank of Japan (DDBJ). A sequence submitted to any one of the three is exchanged and mirrored across all three daily, using shared accession-numbering and format conventions, so a record retrieved from GenBank, ENA, or DDBJ represents the same underlying data regardless of which database a researcher queries.

GenBank itself publishes a new release approximately every two months, with growth statistics published for each release covering both the traditional GenBank divisions and the whole-genome-shotgun (WGS) division separately — an indication that bulk, assembly-scale submissions are tracked as a distinct category from smaller, single-sequence submissions.

GenBank vs. RefSeq: Archival Record vs. Curated Reference

One of the most common points of confusion for researchers and research-support staff new to NCBI’s sequence infrastructure is the difference between GenBank and RefSeq (the Reference Sequence database), and it matters enough for citation and reproducibility purposes to be worth stating precisely:

  • GenBank is an archival, author-submitted database. Records are owned by the submitting lab or sequencing center, are not edited by NCBI beyond basic format validation, and can be redundant — multiple labs may submit essentially the same sequence independently, and older records are not removed when a better one is deposited.
  • RefSeq is a curated, non-redundant reference set derived from GenBank submissions. RefSeq records are owned and maintained by NCBI staff (with varying degrees of manual and computational curation depending on the organism and record type), and are intended to represent a single, stable reference for a given gene, transcript, protein, or genome assembly.

This distinction carries through to genome assemblies specifically: a GenBank assembly (identified with a GCA_ accession prefix) is the archival record as submitted by the originating lab or center, while a RefSeq assembly (GCF_ prefix) is NCBI’s derived copy of that same submission, subject to NCBI’s own annotation pipeline. The two can differ in gene models and annotation even when built from the same underlying assembly, which is a documented, non-trivial source of downstream analysis discrepancies when a project mixes GenBank- and RefSeq-derived annotation without checking which it is using.

Accession Numbers and Versioning

Every GenBank record carries an accession number — a stable identifier for that sequence — and, once revised, an accession-plus-version number (commonly written as accession.version, e.g. an accession followed by .1, .2, and so on) that increments each time the underlying sequence is updated. The base accession number identifies the record’s lineage; the version suffix identifies a specific state of that sequence, which matters for reproducibility: citing an accession without its version leaves open which revision of the sequence a downstream analysis actually used.

How Researchers Submit to GenBank

Submission mechanics differ by scale and record type. For a small number of individual sequences, NCBI’s web-based BankIt tool is the standard route. For larger or more complex submissions — annotated genomes, whole-genome-shotgun (WGS) assemblies, or transcriptome-shotgun-assembly (TSA) datasets — NCBI provides a dedicated Submission Portal and the command-line tool table2asn, which automates creation of the sequence records GenBank requires from a feature table and a set of raw sequences (table2asn is the current tool for this; an older tool, tbl2asn, served a similar purpose for specific submission types such as high-throughput genomic sequence, HTGS, submissions).

One practical point worth flagging for anyone coordinating a submission on behalf of a PI: BankIt assigns a submission (tracking) number at the time of submission, which is not the same as the GenBank accession number and is not suitable for citing in a manuscript. The actual accession number is assigned once NCBI processes the submission — typically within a couple of working days — and that is the identifier that belongs in the paper.

GenBank also supports a confidential hold: sequences submitted ahead of publication can be kept non-public, with the accession number reserved but the record itself withheld, until the associated paper is published — letting a lab satisfy a journal’s pre-submission deposit requirement without releasing the sequence to competitors early.

Why This Matters for Research Administration: Journal Deposit Requirements

Most life-science journals require that any novel DNA, RNA, or amino-acid sequence described in a manuscript be deposited in a public sequence database — GenBank (or an INSDC partner, ENA/DDBJ) for nucleotide data — as a condition of publication, with the resulting accession number cited in the manuscript itself, typically in the methods section or a dedicated data-availability statement. For a research administrator or data steward supporting a life-sciences PI, this is one of the more concrete, deadline-bound intersections between data management and publishing: a manuscript can stall at the review or production stage if the sequence deposit and accession number aren’t in hand when the journal asks for it. Building GenBank (or INSDC) deposit into the same pre-submission checklist as any other data-availability requirement avoids that bottleneck.

GenBank and the Rest of This Cluster: Which Repository for Which Data

Because this cluster already covers several NCBI- and EBI-run genomics repositories, the practical question for anyone deciding where to deposit new data is usually not “is this GenBank-eligible” but “which of these repositories actually matches my data type.” The table below is a quick disambiguation reference; each repository has its own guide with full submission detail.

Repository What it holds Access model
GenBank Raw and assembled DNA/RNA sequences, any organism Open (with optional pre-publication hold)
GEO Gene-expression and functional-genomics experiments Open
dbGaP Individual-level human genotype-phenotype data Controlled access
EGA Individual-level human genomic and phenotypic data (EU) Controlled access
ArrayExpress Functional-genomics/transcriptomics array and sequencing data Open
Metabolomics Workbench Metabolomics experiments and metabolite data (US/NIH) Open
MetaboLights Metabolomics experiments and metabolite data (EU/EBI) Open
ProteomeXchange Mass-spectrometry proteomics data Open
UniProt Curated protein sequence and functional annotation Open
PDB 3D macromolecular structure data Open

A useful rule of thumb: if the object being deposited is a sequence itself, GenBank (or ENA/DDBJ) is very likely the right place. If it’s the output of an experiment performed using sequences — an expression profile, a genotype-phenotype association, a mass-spec proteomics run — one of the specialized repositories above is almost always the correct target, and many of those records will themselves reference GenBank accessions for the underlying sequences involved. GenBank is not a substitute for any of them; it is the substrate several of them build on.

For background on how domain-specific repositories like GenBank fit into repository selection generally, see CASRAI’s guide on how to choose an open data repository and the dictionary entries for discipline-specific repository and trusted digital repository.

Frequently Asked Questions

Is GenBank the same thing as NCBI?

No. NCBI is the agency (part of the National Library of Medicine at NIH) that operates GenBank, along with many other resources — PubMed, GEO, dbGaP, RefSeq, and the broader NCBI Datasets and Entrez systems among them. GenBank specifically refers to the annotated nucleotide sequence archive.

Do I need to submit to GenBank if my sequence is already in GEO or dbGaP?

Usually the underlying sequence itself is what belongs in GenBank (or an INSDC partner); GEO and dbGaP hold the experiment-level data (expression values, genotype-phenotype linkages) that were generated using or associated with that sequence, and typically reference the relevant GenBank/RefSeq accessions rather than duplicating the raw sequence.

How long does it take to get a GenBank accession number?

NCBI generally assigns an accession number within a few working days of a complete, correctly formatted submission, though complex or large submissions (whole genomes, WGS/TSA datasets) can take longer depending on review.

Can I keep a GenBank sequence confidential until my paper is published?

Yes. GenBank supports submitting a sequence ahead of publication and requesting that it remain non-public, with the accession number reserved, until the paper appears — letting authors meet a journal’s pre-submission deposit requirement without releasing the sequence early.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →