Protein Data Bank statistics 2026: size and coverage

Matic BrozComputational chemist
The Protein Data Bank is the global open archive for experimentally determined macromolecular structures. As of August 9, 2026, it contained 258,023 archive entries. Of these, 252,619 contained protein and 221,278 were protein-only entries.
RCSB PDB also indexed 1,062,058 computed structure models on the same date. Those predictions are searchable alongside the archive, but they are not experimental PDB entries.
How large is the Protein Data Bank?
The Protein Data Bank contained 258,023 current archive entries on August 9, 2026, and 252,619 of them contained at least one protein.
| Count | Entries or models | What it measures |
|---|---|---|
| Entire PDB archive | 258,023 | Experimental and integrative archive entries |
| Entries containing protein | 252,619 | Protein-only, protein/nucleic-acid, and protein/oligosaccharide entries |
| Protein-only entries | 221,278 | Entries whose polymer content is protein only |
| Computed models on RCSB.org | 1,062,058 | Predictions indexed separately from the PDB archive |
RCSB reports the total, protein-only, protein/nucleic-acid, and protein/oligosaccharide categories directly. The 252,619 protein-containing total is calculated as 221,278 + 16,783 + 14,558, or 97.9% of the archive.[1] Computed models are a separate RCSB.org collection.[9]
A PDB entry is not a unique protein. One protein can have many entries for different ligands, mutations, conformations, complexes, or experimental conditions. An NMR entry can also contain several coordinate models. For that reason, archive entries, protein sequences, biological proteins, coordinate models, and computed predictions should not be added together.
How many unique protein sequences are represented in the PDB?
The current RCSB sequence-cluster files contain 135,501 distinct 100%-identity clusters with at least one PDB protein polymer entity.
| Sequence identity threshold | Clusters containing a PDB entity |
|---|---|
| 100% | 135,501 |
| 90% | 93,735 |
| 50% | 58,556 |
RCSB clusters protein polymer entities weekly with DIAMOND and publishes one cluster per line.[2] ProteinIQ counted lines containing at least one four-character PDB entity ID in the August 4 files, excluding clusters made up only of computed-model IDs. The source files contained 135,501 such lines at 100% identity, 93,735 at 90%, and 58,556 at 50%.[3][4][5]
These are sequence clusters, not counts of unique biological proteins. Engineered constructs, fragments, mutations, and identical sequences from different organisms affect the result. A lower identity threshold merges more distant sequences into fewer clusters.
There is also no universal percentage for protein structure coverage. Coverage changes with the chosen species, reference proteome, sequence-identity threshold, and whether partial domains count. The broader protein count therefore cannot be compared directly with the PDB entry total.
Which experimental methods contribute most PDB entries?
X-ray crystallography accounts for 206,356 current PDB entries, or 80.0% of the archive. Electron microscopy contributes 36,069 entries (14.0%), NMR contributes 14,811 (5.7%), and all other method categories together contribute 787 (0.3%).[1]
The percentages are calculated from the 258,023-entry total and rounded to one decimal place. “Other methods” combines 394 integrative entries, 266 multiple-method entries, 90 neutron entries, and 37 entries in RCSB's other category.[1]
Experimental method affects which comparisons are sensible. FoldSeek searches structures by shape, while tools such as US-align compare selected structures in detail. Neither removes the need to check method, resolution, construct, and biological assembly.
How quickly is the PDB growing?
The PDB released 17,610 structures during 2025, its largest complete release year in the current RCSB growth table. A further 10,787 structures were released in 2026 through August 9, bringing the archive to 258,023 entries.[6]
Depositions and releases are different counts. wwPDB recorded 20,975 deposits in 2025 and 13,279 in 2026 through August 4. Its deposition series can include entries later withdrawn or obsoleted, and a deposit may be released in a different year.[7]
The archive also recorded 4.72 billion coordinate, experimental-data, validation-report, and website download or view events in 2025. The 2026 year-to-date total reached 3.26 billion on August 3.[8] These figures measure access events, not unique people or unique structures.
Are computed structure models part of the PDB?
Computed structure models are not part of the PDB archive. RCSB.org indexed 1,062,058 of them on August 9, 2026, about 4.1 models for every current PDB entry.[1]
RCSB integrates selected models from AlphaFold DB and ModelArchive so users can search experimental structures and predictions through one interface.[9] The AlphaFold statistics page tracks the much larger AlphaFold database itself; its totals should not be used as PDB archive counts.
Experimental entries and predictions also carry different evidence. A PDB entry reports a deposited structural experiment or integrative determination. A computed model reports a prediction whose confidence can vary by residue and whose biological assembly may be unknown.
How should PDB structure data be downloaded?
PDBx/mmCIF has been the standard PDB archive distribution format since 2014. The legacy .pdb format is frozen and cannot fully represent entries with more than 62 chains or more than 99,999 ATOM records.[10]
ProteinIQ's PDB downloader retrieves archive structures by accession, and the PDB-to-CIF converter converts legacy coordinate files when a workflow requires mmCIF.
wwPDB will stop issuing four-character accessions on July 21, 2027, when it moves fully to extended 12-character PDB IDs. Existing IDs will gain a padded form, such as 1abc becoming pdb_00001abc.[11] PDB archive data files are released under the CC0 1.0 Universal Public Domain Dedication.[12]
Sources▼
- PDB data distribution by experimental method and molecular type RCSB PDB · August 9, 2026. https://www.rcsb.org/stats/summary
- File download services: sequence clusters data RCSB PDB · August 9, 2026. https://www.rcsb.org/docs/programmatic-access/file-download-services
- Protein sequence clusters at 100% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-100.txt
- Protein sequence clusters at 90% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-90.txt
- Protein sequence clusters at 50% identity RCSB PDB · August 9, 2026. https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-50.txt
- Overall growth of released PDB structures per year RCSB PDB · August 9, 2026. https://www.rcsb.org/stats/growth/growth-released-structures
- PDB deposition statistics wwPDB · August 9, 2026. https://www.wwpdb.org/stats/deposition
- PDB download statistics wwPDB · August 9, 2026. https://www.wwpdb.org/stats/download
- Computed Structure Models and RCSB.org RCSB PDB · August 9, 2026. https://www.rcsb.org/docs/general-help/computed-structure-models-and-rcsborg
- File formats and the PDB wwPDB · August 9, 2026. https://www.wwpdb.org/documentation/file-formats-and-the-pdb
- Extended PDB ID with 12 characters wwPDB · August 9, 2026. https://www.wwpdb.org/documentation/new-format-for-pdb-ids
- Usage policies wwPDB · August 9, 2026. https://www.wwpdb.org/about/usage-policies

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.