ProteinIQ
Sign inStart for free
ProteinIQ
Statistics/Aug '26/5 min read

Protein Data Bank statistics 2026: size and coverage

Matic Broz

Matic BrozComputational chemist

The Protein Data Bank is the global open archive for experimentally determined macromolecular structures. As of August 9, 2026, it contained 258,023 archive entries. Of these, 252,619 contained protein and 221,278 were protein-only entries.

RCSB PDB also indexed 1,062,058 computed structure models on the same date. Those predictions are searchable alongside the archive, but they are not experimental PDB entries.

How large is the Protein Data Bank?

The Protein Data Bank contained 258,023 current archive entries on August 9, 2026, and 252,619 of them contained at least one protein.

CountEntries or modelsWhat it measures
Entire PDB archive258,023Experimental and integrative archive entries
Entries containing protein252,619Protein-only, protein/nucleic-acid, and protein/oligosaccharide entries
Protein-only entries221,278Entries whose polymer content is protein only
Computed models on RCSB.org1,062,058Predictions indexed separately from the PDB archive

RCSB reports the total, protein-only, protein/nucleic-acid, and protein/oligosaccharide categories directly. The 252,619 protein-containing total is calculated as 221,278 + 16,783 + 14,558, or 97.9% of the archive.[1] Computed models are a separate RCSB.org collection.[9]

A PDB entry is not a unique protein. One protein can have many entries for different ligands, mutations, conformations, complexes, or experimental conditions. An NMR entry can also contain several coordinate models. For that reason, archive entries, protein sequences, biological proteins, coordinate models, and computed predictions should not be added together.

How many unique protein sequences are represented in the PDB?

The current RCSB sequence-cluster files contain 135,501 distinct 100%-identity clusters with at least one PDB protein polymer entity.

Sequence identity thresholdClusters containing a PDB entity
100%135,501
90%93,735
50%58,556

RCSB clusters protein polymer entities weekly with DIAMOND and publishes one cluster per line.[2] ProteinIQ counted lines containing at least one four-character PDB entity ID in the August 4 files, excluding clusters made up only of computed-model IDs. The source files contained 135,501 such lines at 100% identity, 93,735 at 90%, and 58,556 at 50%.[3][4][5]

These are sequence clusters, not counts of unique biological proteins. Engineered constructs, fragments, mutations, and identical sequences from different organisms affect the result. A lower identity threshold merges more distant sequences into fewer clusters.

There is also no universal percentage for protein structure coverage. Coverage changes with the chosen species, reference proteome, sequence-identity threshold, and whether partial domains count. The broader protein count therefore cannot be compared directly with the PDB entry total.

Which experimental methods contribute most PDB entries?

X-ray crystallography accounts for 206,356 current PDB entries, or 80.0% of the archive. Electron microscopy contributes 36,069 entries (14.0%), NMR contributes 14,811 (5.7%), and all other method categories together contribute 787 (0.3%).[1]

Current PDB entries by experimental method: 80.0% X-ray, 14.0% electron microscopy, 5.7% NMR, and 0.3% other methods

The percentages are calculated from the 258,023-entry total and rounded to one decimal place. “Other methods” combines 394 integrative entries, 266 multiple-method entries, 90 neutron entries, and 37 entries in RCSB's other category.[1]

Experimental method affects which comparisons are sensible. FoldSeek searches structures by shape, while tools such as US-align compare selected structures in detail. Neither removes the need to check method, resolution, construct, and biological assembly.

How quickly is the PDB growing?

The PDB released 17,610 structures during 2025, its largest complete release year in the current RCSB growth table. A further 10,787 structures were released in 2026 through August 9, bringing the archive to 258,023 entries.[6]

Depositions and releases are different counts. wwPDB recorded 20,975 deposits in 2025 and 13,279 in 2026 through August 4. Its deposition series can include entries later withdrawn or obsoleted, and a deposit may be released in a different year.[7]

The archive also recorded 4.72 billion coordinate, experimental-data, validation-report, and website download or view events in 2025. The 2026 year-to-date total reached 3.26 billion on August 3.[8] These figures measure access events, not unique people or unique structures.

Are computed structure models part of the PDB?

Computed structure models are not part of the PDB archive. RCSB.org indexed 1,062,058 of them on August 9, 2026, about 4.1 models for every current PDB entry.[1]

RCSB integrates selected models from AlphaFold DB and ModelArchive so users can search experimental structures and predictions through one interface.[9] The AlphaFold statistics page tracks the much larger AlphaFold database itself; its totals should not be used as PDB archive counts.

Experimental entries and predictions also carry different evidence. A PDB entry reports a deposited structural experiment or integrative determination. A computed model reports a prediction whose confidence can vary by residue and whose biological assembly may be unknown.

How should PDB structure data be downloaded?

PDBx/mmCIF has been the standard PDB archive distribution format since 2014. The legacy .pdb format is frozen and cannot fully represent entries with more than 62 chains or more than 99,999 ATOM records.[10]

ProteinIQ's PDB downloader retrieves archive structures by accession, and the PDB-to-CIF converter converts legacy coordinate files when a workflow requires mmCIF.

wwPDB will stop issuing four-character accessions on July 21, 2027, when it moves fully to extended 12-character PDB IDs. Existing IDs will gain a padded form, such as 1abc becoming pdb_00001abc.[11] PDB archive data files are released under the CC0 1.0 Universal Public Domain Dedication.[12]

Sources▼
  1. PDB data distribution by experimental method and molecular type RCSB PDB · August 9, 2026. https://www.rcsb.org/stats/summary
  2. File download services: sequence clusters data RCSB PDB · August 9, 2026. https://www.rcsb.org/docs/programmatic-access/file-download-services
  3. Protein sequence clusters at 100% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-100.txt
  4. Protein sequence clusters at 90% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-90.txt
  5. Protein sequence clusters at 50% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-50.txt
  6. Overall growth of released PDB structures per year RCSB PDB · August 9, 2026. https://www.rcsb.org/stats/growth/growth-released-structures
  7. PDB deposition statistics wwPDB · August 9, 2026. https://www.wwpdb.org/stats/deposition
  8. PDB download statistics wwPDB · August 9, 2026. https://www.wwpdb.org/stats/download
  9. Computed Structure Models and RCSB.org RCSB PDB · August 9, 2026. https://www.rcsb.org/docs/general-help/computed-structure-models-and-rcsborg
  10. File formats and the PDB wwPDB · August 9, 2026. https://www.wwpdb.org/documentation/file-formats-and-the-pdb
  11. Extended PDB ID with 12 characters wwPDB · August 9, 2026. https://www.wwpdb.org/documentation/new-format-for-pdb-ids
  12. Usage policies wwPDB · August 9, 2026. https://www.wwpdb.org/about/usage-policies
Published
June 30, 2026
Last updated
August 9, 2026

Table of contents

Cite this article

Broz, M. (2026, August 9). Protein Data Bank statistics 2026: size and coverage. ProteinIQ. https://proteiniq.io/guides/rcsb-statistics

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

Related guides

Statistics

AUG '26

AlphaFold database statistics [2026]

AlphaFold DB contains 262,739,159 predicted models as of August 9, 2026. See its release history, human-proteome coverage, confidence thresholds, and relationship to the PDB.

Matic Broz Computational chemist

Statistics

AUG '26

How many types of antibodies are there?

The human body has five antibody classes. Compare the structures, locations, functions, and subclasses of IgG, IgA, IgM, IgD, and IgE.

Matic Broz Computational chemist

Statistics

AUG '26

How accurate is DNA sequencing?

DNA sequencing accuracy ranges from about 99% per raw base to above 99.9% for high-quality or consensus reads. The exact figure depends on the platform, chemistry, software, sample, and metric.

Matic Broz Computational chemist

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing