ProteinIQ
Sign inStart for free
ProteinIQ
Genetics

How many genomes have been sequenced?

NCBI indexes 3.67 million current public genome assemblies, including 3.25 million bacterial assemblies. GenBank release 272 contains 6.52 billion sequence records.

August 10, 2026·Matic Broz, PhD
DNA sequence records arranged as an engraved scientific archive.

NCBI indexes 3,667,228 current public genome assemblies as of August 10, 2026. This is the best reproducible count of genomes sequenced and assembled in a public archive. It counts assembly records rather than unique species, people, or sequencing runs.

Most of these records are bacterial. GenBank itself contains a much larger 6.52 billion sequence records because a genome assembly can consist of many sequence records, and many records contain only part of a genome.

How many genomes have been sequenced?

NCBI Datasets indexes 3,667,228 current public genome assemblies as of August 10, 2026.[4][5]

The count excludes paired reports so that a genome represented in both GenBank and RefSeq appears once. It includes assemblies at contig, scaffold, chromosome, and complete-genome level. A count limited to fully sequenced genomes would therefore be smaller.[4][5]

The total leaves out genomes that were sequenced but never deposited as current public assemblies. Raw-read experiments and genome assemblies are separate archive units.[3][4] The 3.67 million figure answers a narrower question that can be checked and repeated: how many current public assembly records does NCBI index?

How many bacterial genomes have been sequenced?

NCBI lists 3,251,345 current bacterial genome assemblies, or 88.7% of its all-taxa total under the same filter.[6]

The 3.25 million figure counts assemblies rather than bacterial species. One species, strain, or isolate may have many assemblies, and an assembly may be a draft made of contigs or scaffolds rather than a finished chromosome.

The scale reflects repeated sequencing across clinical surveillance, food safety, environmental sampling, and strain comparison. MUMmer4 can compare related assemblies and locate substitutions, insertions, deletions, and larger structural differences.

How many plant genomes have been sequenced?

NCBI lists 9,558 current plant genome assemblies as of August 10, 2026.[8] The total includes multiple cultivars, individuals, and assembly versions for some species.

The NCBI assembly total cannot answer how many human genomes have been sequenced worldwide. NCBI lists 2,574 current human assemblies under the same filter, but this counts public assemblies rather than every person whose DNA has gone through whole-genome sequencing.[9] Across all eukaryotes, NCBI lists 67,257 current assemblies.[7]

NCBI Datasets current genome assembly counts for bacteria, eukaryotes, plants, and humans on a logarithmic scale

Plants and humans are subsets of eukaryotes, so the four chart categories overlap rather than forming parts of one total.

There is no direct answer to how many species' genomes have been sequenced in these assembly totals. A species can have many assemblies, and quality thresholds change which draft genomes qualify. The latest peer-reviewed GenBank update reported nucleotide sequences from more than 581,000 formally described species, but a single gene or locus is enough for a species to enter that count.[3] The figure therefore includes species without whole genomes.

How many sequences and bases are in GenBank?

GenBank release 272 contains 6.52 billion sequence records and 57.69 trillion nucleotide bases.[1]

Record groupSequence recordsNucleotide bases
Traditional GenBank264,214,3547,618,210,921,117
Set-based WGS, TSA, and TLS6,255,423,20450,068,290,456,326
Total6,519,637,55857,686,501,377,443

The figures are the traditional and set-based totals reported in the GenBank 272.0 distribution notes.[1] Set-based records account for 95.9% of all records and 86.8% of all bases.

WGS records contain whole-genome shotgun assemblies, TSA records contain transcriptome shotgun assemblies, and TLS records contain targeted locus studies. Raw next-generation sequencing reads are generally stored in the Sequence Read Archive; GenBank holds assembled nucleotide sequences and their annotations.[3]

A GenBank record therefore contains more than a bare sequence. Converting a downloaded GenBank file to FASTA keeps the selected sequence while leaving out most of that annotation structure.

How fast has GenBank grown?

In NCBI's historical GenBank and WGS series, the amount of sequence data grew from 680,338 bases in December 1982 to 56.70 trillion bases in June 2026.[2]

Combined traditional GenBank and whole-genome shotgun sequence data grew from 680,338 bases in 1982 to 56.70 trillion bases in 2026

That is an 83.3-million-fold increase. Over the same period, the combined number of traditional GenBank and WGS records rose from 606 to 5,258,544,876, an 8.68-million-fold increase.[2]

The chart follows NCBI's published GenBank-and-WGS series so that every point uses the same definition. Its 2026 endpoint is lower than the full 57.69-trillion-base release total because the full release notes also include TSA and TLS set-based records.[1][2]

Sources9 references
  1. Current GenBank release notes: release 272.0

    National Center for Biotechnology Information · August 10, 2026

  2. GenBank and WGS statistics

    National Center for Biotechnology Information · August 10, 2026

  3. GenBank 2025 update

    Nucleic Acids Research · 2025

  4. NCBI Datasets genome API filters

    National Center for Biotechnology Information · August 10, 2026

  5. NCBI Datasets current genome assembly count for all taxa, excluding paired reports

    National Center for Biotechnology Information · August 10, 2026

  6. NCBI Datasets current bacterial genome assembly count, excluding paired reports

    National Center for Biotechnology Information · August 10, 2026

  7. NCBI Datasets current eukaryotic genome assembly count, excluding paired reports

    National Center for Biotechnology Information · August 10, 2026

  8. NCBI Datasets current plant genome assembly count, excluding paired reports

    National Center for Biotechnology Information · August 10, 2026

  9. NCBI Datasets current human genome assembly count, excluding paired reports

    National Center for Biotechnology Information · August 10, 2026

About the author

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

LinkedInGoogle ScholarORCID
Published
July 30, 2026
Last updated
August 10, 2026

Related guides

Browse all guides
A magnifying glass examines DNA alongside overlapping sequence fragments.

Genetics · August 10, 2026

How accurate is DNA sequencing?

DNA sequencing accuracy ranges from about 99% per raw base to above 99.9% for high-quality or consensus reads. The exact figure depends on the platform, chemistry, software, sample, and metric.

Cutaway mitochondrion with an enlarged circular mitochondrial DNA molecule.

Genetics · September 19, 2026

How many genes are in mitochondrial DNA?

Human mitochondrial DNA has 37 canonical genes. Compare the standard count with evidence for additional small proteins and mtDNA copy numbers across tissues.

Illustration comparing a human, chimpanzee, mouse, dog, chicken, and zebrafish.

Genetics · September 19, 2026

How much DNA do humans share with other animals?

Human and chimpanzee DNA is about 98.8% identical across directly aligned bases. See how animal comparisons change when studies count DNA letters, proteins, or shared genes.

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • Batches
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation
  • For academia
  • For enterprise

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing