ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Structures

Protein Data Bank statistics 2026: size and coverage

August 9, 2026·Matic Broz, PhD
Ink illustration of archived protein structure records and a crystal representing the Protein Data Bank.

The Protein Data Bank is the global open archive for experimentally determined macromolecular structures. As of August 9, 2026, it contained 258,023 archive entries. Of these, 252,619 contained protein and 221,278 were protein-only entries.

RCSB PDB also indexed 1,062,058 computed structure models on the same date. Those predictions are searchable alongside the archive, but they are not experimental PDB entries.

How large is the Protein Data Bank?

The Protein Data Bank contained 258,023 current archive entries on August 9, 2026, and 252,619 of them contained at least one protein.

CountEntries or modelsWhat it measures
Entire PDB archive258,023Experimental and integrative archive entries
Entries containing protein252,619Protein-only, protein/nucleic-acid, and protein/oligosaccharide entries
Protein-only entries221,278Entries whose polymer content is protein only
Computed models on RCSB.org1,062,058Predictions indexed separately from the PDB archive

RCSB reports the total, protein-only, protein/nucleic-acid, and protein/oligosaccharide categories directly. The 252,619 protein-containing total is calculated as 221,278 + 16,783 + 14,558, or 97.9% of the archive.[1] Computed models are a separate RCSB.org collection.[9]

A PDB entry is not a unique protein. One protein can have many entries for different ligands, mutations, conformations, complexes, or experimental conditions. An NMR entry can also contain several coordinate models. For that reason, archive entries, protein sequences, biological proteins, coordinate models, and computed predictions should not be added together.

How many unique protein sequences are represented in the PDB?

The current RCSB sequence-cluster files contain 135,501 distinct 100%-identity clusters with at least one PDB protein polymer entity.

Sequence identity thresholdClusters containing a PDB entity
100%135,501
90%93,735
50%58,556

RCSB clusters protein polymer entities weekly with DIAMOND and publishes one cluster per line.[2] ProteinIQ counted lines containing at least one four-character PDB entity ID in the August 4 files, excluding clusters made up only of computed-model IDs. The source files contained 135,501 such lines at 100% identity, 93,735 at 90%, and 58,556 at 50%.[3][4][5]

These are sequence clusters, not counts of unique biological proteins. Engineered constructs, fragments, mutations, and identical sequences from different organisms affect the result. A lower identity threshold merges more distant sequences into fewer clusters.

There is also no universal percentage for protein structure coverage. Coverage changes with the chosen species, reference proteome, sequence-identity threshold, and whether partial domains count. The broader protein count therefore cannot be compared directly with the PDB entry total.

Which experimental methods contribute most PDB entries?

X-ray crystallography accounts for 206,356 current PDB entries, or 80.0% of the archive. Electron microscopy contributes 36,069 entries (14.0%), NMR contributes 14,811 (5.7%), and all other method categories together contribute 787 (0.3%).[1]

Current PDB entries by experimental method: 80.0% X-ray, 14.0% electron microscopy, 5.7% NMR, and 0.3% other methods. Reuse under CC BY 4.0.

The percentages are calculated from the 258,023-entry total and rounded to one decimal place. “Other methods” combines 394 integrative entries, 266 multiple-method entries, 90 neutron entries, and 37 entries in RCSB's other category.[1]

Experimental method affects which comparisons are sensible. FoldSeek searches structures by shape, while tools such as US-align compare selected structures in detail. Neither removes the need to check method, resolution, construct, and biological assembly.

How quickly is the PDB growing?

The PDB released 17,610 structures during 2025, its largest complete release year in the current RCSB growth table. A further 10,787 structures were released in 2026 through August 9, bringing the archive to 258,023 entries.[6]

Depositions and releases are different counts. wwPDB recorded 20,975 deposits in 2025 and 13,279 in 2026 through August 4. Its deposition series can include entries later withdrawn or obsoleted, and a deposit may be released in a different year.[7]

The archive also recorded 4.72 billion coordinate, experimental-data, validation-report, and website download or view events in 2025. The 2026 year-to-date total reached 3.26 billion on August 3.[8] These figures measure access events, not unique people or unique structures.

Are computed structure models part of the PDB?

Computed structure models are not part of the PDB archive. RCSB.org indexed 1,062,058 of them on August 9, 2026, about 4.1 models for every current PDB entry.[1]

RCSB integrates selected models from AlphaFold DB and ModelArchive so users can search experimental structures and predictions through one interface.[9] The AlphaFold statistics page tracks the much larger AlphaFold database itself; its totals should not be used as PDB archive counts.

Experimental entries and predictions also carry different evidence. A PDB entry reports a deposited structural experiment or integrative determination. A computed model reports a prediction whose confidence can vary by residue and whose biological assembly may be unknown.

How should PDB structure data be downloaded?

PDBx/mmCIF has been the standard PDB archive distribution format since 2014. The legacy .pdb format is frozen and cannot fully represent entries with more than 62 chains or more than 99,999 ATOM records.[10]

ProteinIQ's PDB downloader retrieves archive structures by accession, and the PDB-to-CIF converter converts legacy coordinate files when a workflow requires mmCIF.

wwPDB will stop issuing four-character accessions on July 21, 2027, when it moves fully to extended 12-character PDB IDs. Existing IDs will gain a padded form, such as 1abc becoming pdb_00001abc.[11] PDB archive data files are released under the CC0 1.0 Universal Public Domain Dedication.[12]

Sources12
  1. PDB data distribution by experimental method and molecular type

    RCSB PDB · August 9, 2026

  2. File download services: sequence clusters data

    RCSB PDB · August 9, 2026

  3. Protein sequence clusters at 100% identity

    RCSB PDB · August 9, 2026

  4. Protein sequence clusters at 90% identity

    RCSB PDB · August 9, 2026

  5. Protein sequence clusters at 50% identity

    RCSB PDB · August 9, 2026

  6. Overall growth of released PDB structures per year

    RCSB PDB · August 9, 2026

  7. PDB deposition statistics

    wwPDB · August 9, 2026

  8. PDB download statistics

    wwPDB · August 9, 2026

  9. Computed Structure Models and RCSB.org

    RCSB PDB · August 9, 2026

  10. File formats and the PDB

    wwPDB · August 9, 2026

  11. Extended PDB ID with 12 characters

    wwPDB · August 9, 2026

  12. Usage policies

    wwPDB · August 9, 2026

Cite this article

Broz, M. (2026, August 9). Protein Data Bank statistics 2026: size and coverage. ProteinIQ. https://proteiniq.io/guides/rcsb-statistics

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “Protein Data Bank statistics 2026: size and coverage” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
June 30, 2026
Updated
August 9, 2026

Related guides

Browse all guides
Engraved protein folds illustrating a collection of predicted structures.

Structures · September 24, 2026

AlphaFold database statistics [2026]

AlphaFold DB lists 261,552,403 predicted models as of September 24, 2026. See the collection breakdown, release timeline, human proteome coverage, pLDDT, PAE and interface thresholds, and how the database compares with the PDB.

Two protein chains with low-PAE residue pairs highlighted at their interface and faint flexible tails

Structures · October 2, 2026

ipSAE explained

ipSAE is an interface confidence score for AlphaFold2, AlphaFold3 and Boltz predictions that counts only confidently placed residue pairs between chains. Learn how it is calculated, why it beats ipTM on full-length sequences, what 0.6, 0.7 and 0.8 mean, and how to calculate it for your own complexes and binder designs.

Two-chain protein complex with pTM marking the overall structure and ipTM marking relationships between chains

Structures · October 1, 2026

ipTM vs pTM: judging predicted complexes

pTM scores the whole predicted structure and ipTM scores only the placement of chains relative to each other. Learn how both are calculated, what the 0.5, 0.6 and 0.8 cutoffs mean, why disordered tails lower ipTM, and how to read them in AlphaFold2, Boltz-2 and other complex predictors.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing