3 min read
How many sequences are in GenBank?
GenBank contains 6.52 billion sequence records and 57.69 trillion nucleotide bases in release 272. See how the database has grown and how many public bacterial, plant, and human genome assemblies NCBI indexes.

Matic Broz Computational chemist
GenBank contains 6,519,637,558 sequence records and 57,686,501,377,443 nucleotide bases in release 272, published on June 15, 2026. Rounded, that is 6.52 billion records and 57.69 trillion bases.
Each total counts records in a public nucleotide archive. Genomes, species, and people are different units, and their counts are much lower.
How many sequences and bases are in GenBank?
GenBank release 272 contains 6.52 billion sequence records and 57.69 trillion nucleotide bases.[1]
| Record group | Sequence records | Nucleotide bases |
|---|---|---|
| Traditional GenBank | 264,214,354 | 7,618,210,921,117 |
| Set-based WGS, TSA, and TLS | 6,255,423,204 | 50,068,290,456,326 |
| Total | 6,519,637,558 | 57,686,501,377,443 |
The figures are the traditional and set-based totals reported in the GenBank 272.0 distribution notes.[1] Set-based records account for 95.9% of all records and 86.8% of all bases.
WGS records contain whole-genome shotgun assemblies, TSA records contain transcriptome shotgun assemblies, and TLS records contain targeted locus studies. Raw next-generation sequencing reads are generally stored in the Sequence Read Archive; GenBank holds assembled nucleotide sequences and their annotations.[3]
A GenBank record therefore contains more than a bare sequence. Converting a downloaded GenBank file to FASTA keeps the selected sequence while leaving out most of that annotation structure.
How fast has GenBank grown?
In NCBI's historical GenBank and WGS series, the amount of sequence data grew from 680,338 bases in December 1982 to 56.70 trillion bases in June 2026.[2]
That is an 83.3-million-fold increase. Over the same period, the combined number of traditional GenBank and WGS records rose from 606 to 5,258,544,876, an 8.68-million-fold increase.[2]
The chart follows NCBI's published GenBank-and-WGS series so that every point uses the same definition. Its 2026 endpoint is lower than the full 57.69-trillion-base release total because the full release notes also include TSA and TLS set-based records.[1][2]
How many genomes have been sequenced?
NCBI Datasets indexed 3,651,813 current public genome assemblies on July 30, 2026, with one report returned for each paired GenBank and RefSeq assembly.[4][5]
This reproducible count measures public assembly records under a defined filter. Genomes held privately or only as raw reads fall outside it, and the same organism can have many assemblies. One assembly can also contain hundreds or thousands of sequence records.
The number of species with sequence data is larger than the number with assembled genomes. The most recent peer-reviewed GenBank update reported sequences from more than 581,000 formally described species in 2025.[3] A species can enter that count through a single gene or locus. The 581,000 figure therefore includes species without a whole-genome assembly.
The distinction is large even for familiar organisms. A human genome contains roughly 3.1 billion base pairs, while an individual GenBank record may contain one chromosome, contig, transcript, gene, or targeted region.
How many bacterial, plant, and human genomes have been sequenced?
As of July 30, 2026, NCBI Datasets listed 3,236,975 bacterial assemblies, 66,783 eukaryotic assemblies, 9,394 plant assemblies, and 2,572 human assemblies after excluding paired GenBank and RefSeq reports.[6][7][8][9]
Plants and humans are subsets of eukaryotes, so the four bars show selected taxonomic scopes rather than parts of one total. Bacteria account for 88.6% of the 3,651,813 all-taxa assembly count under the same filter.
The 2,572 human entries measure public assembly records rather than the global number of people whose genomes have been sequenced. The 9,394 plant entries include multiple cultivars, individuals, and assembly versions for some species.
When two assemblies represent related organisms or strains, MUMmer4 can compare them at nucleotide level and locate substitutions, insertions, deletions, and larger structural differences.
Sources▼
- Current GenBank release notes: release 272.0 National Center for Biotechnology Information · July 30, 2026. https://www.ncbi.nlm.nih.gov/genbank/release/current/
- GenBank and WGS statistics National Center for Biotechnology Information · July 30, 2026. https://www.ncbi.nlm.nih.gov/genbank/statistics/
- GenBank 2025 update Nucleic Acids Research · 2025. https://academic.oup.com/nar/article/53/D1/D56/7903376
- NCBI Datasets genome API filters National Center for Biotechnology Information · July 30, 2026. https://www.ncbi.nlm.nih.gov/datasets/docs/v2/api/rest-api/
- NCBI Datasets current genome assembly count for all taxa, excluding paired reports National Center for Biotechnology Information · July 30, 2026. https://api.ncbi.nlm.nih.gov/datasets/v2/genome/taxon/1/dataset_report?page_size=1&filters.exclude_paired_reports=true
- NCBI Datasets current bacterial genome assembly count, excluding paired reports National Center for Biotechnology Information · July 30, 2026. https://api.ncbi.nlm.nih.gov/datasets/v2/genome/taxon/2/dataset_report?page_size=1&filters.exclude_paired_reports=true
- NCBI Datasets current eukaryotic genome assembly count, excluding paired reports National Center for Biotechnology Information · July 30, 2026. https://api.ncbi.nlm.nih.gov/datasets/v2/genome/taxon/2759/dataset_report?page_size=1&filters.exclude_paired_reports=true
- NCBI Datasets current plant genome assembly count, excluding paired reports National Center for Biotechnology Information · July 30, 2026. https://api.ncbi.nlm.nih.gov/datasets/v2/genome/taxon/33090/dataset_report?page_size=1&filters.exclude_paired_reports=true
- NCBI Datasets current human genome assembly count, excluding paired reports National Center for Biotechnology Information · July 30, 2026. https://api.ncbi.nlm.nih.gov/datasets/v2/genome/taxon/9606/dataset_report?page_size=1&filters.exclude_paired_reports=true

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.