ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Structures

AlphaFold database statistics [2026]

September 24, 2026·Matic Broz, PhD
Engraved protein folds illustrating a collection of predicted structures.

The AlphaFold Protein Structure Database lists 261,552,403 predicted protein models as of September 24, 2026. About 241 million of them form the core release, one AlphaFold 2 prediction for most sequences in UniProt, the central catalogue of known proteins. The rest come from partner datasets that add bacterial genomes, viruses, parasites, and predicted protein pairs.

That figure counts computer predictions, not experimentally determined structures, and one protein can have several entries. For scale, the Protein Data Bank holds about 260,000 experimental entries, so AlphaFold DB offers roughly a thousand predictions for every structure solved in a laboratory.

How many structures are in the AlphaFold Database?

The AlphaFold DB FAQ lists 261,552,403 predicted models, including 40,054 isoforms, with 46 complete proteomes available for bulk download.[1] The latest numbered core release, v6 from October 2025, contains 241,070,489 predictions synchronized to UniProt release 2025_03.[2] The difference between the two numbers is the partner collections added since then.

CollectionPredicted modelsWhat it covers
UniProt core release (v6)241,070,489Single chains for UniProt 2025_03, including 40,054 isoforms
AllTheBacteriaover 17.7 millionHypothetical proteins from more than 2.4 million bacterial and archaeal genomes
NVIDIA proteome-scale dimers2,158,419 homodimers and 79,156 heterodimersProtein pairs from Swiss-Prot, 16 model organisms, and 30 WHO global health proteomes
Viral complexes1,703,992Pairwise viral protein interactions across about 2,800 viral proteomes
Big Fantastic Virus Database (BFVD)351,242Viral protein sequences from UniRef30 2023_02
Viro3D82,424Proteins from 4,407 human and animal viruses
Viral AlphaFold Database (VAD)53,716Full-length viral RefSeq proteins, as monomers and homodimers
Kinetoplastids16,887Trypanosoma and Leishmania proteins
Table 1. What the AlphaFold DB total is made of. Collection sizes as reported by AlphaFold DB on September 24, 2026. For the dimer datasets, the counts are the models shown on entry pages after filtering by interface confidence (pDockQ2 of at least 0.23 and ipSAE of at least 0.6). Sources: AlphaFold DB FAQ (2026), PDBe release notes (2025).[1][2]

By our tally, the collection sizes excluding the viral complexes add up to about 261.5 million, within 0.02% of the headline figure. The viral-complex release arrived in September 2026 and appears not yet to be included in the FAQ counter, although the FAQ does not say which collections its total covers.[1] The count is also revised downward at times. In August 2026 the same FAQ listed 262,739,159 models, and an EMBL-EBI training page still showed 260,986,406 on September 24, 2026.[3] Anyone citing a single number should give the page and the access date.

These totals count database entries, not distinct proteins. One UniProt accession can have a canonical model, isoform models (a hyphenated suffix such as P42167-2), fragments of a long protein (such as Q8WZ42-F1, a piece of titin, the largest human protein), and separate predictions from partner datasets made with different software.[1]

Figure 1. Growth of the AlphaFold Database. Points mark the launch, the v4 and v6 core releases, and the website total on September 24, 2026, and are not evenly spaced in time. Sources: Varadi et al. (2024), PDBe release notes, AlphaFold DB FAQ. Reuse under CC BY 4.0.

AlphaFold DB launched in July 2021 with more than 360,000 structures from 20 model-organism proteomes, and version 4 contained 214,683,829 predictions.[4][2] Nearly all of that growth happened in a single step in July 2022, when most of UniProt was added at once. Since then the core has grown slowly with UniProt, and most new entries come from partner datasets. The full data can be retrieved from the EBI FTP site and Google Cloud; the v4 Google Cloud dataset alone was about 23 TiB.[4] Individual model files can also be fetched with the AlphaFold database downloader.

When was AlphaFold released, and how has the database changed?

AlphaFold 2 was presented at the CASP14 prediction assessment on November 30, 2020, and its method paper and source code were published on July 15, 2021. The database followed a week later, on July 22, 2021.[5][6] The database still contains AlphaFold 2 and related ColabFold predictions. AlphaFold 3, released with the AlphaFold Server on May 8, 2024, is not the source of its models.[5][7]

DateEvent
November 30, 2020AlphaFold 2 tops the CASP14 assessment
July 15, 2021AlphaFold 2 paper published in Nature and code released
July 22, 2021AlphaFold DB launches with 20 model-organism proteomes, more than 360,000 structures
December 2021Swiss-Prot added
January 2022WHO global health proteomes added
July 28, 2022Most of UniProt added, taking the database past 200 million structures
November 2022v4 fixes a numerical bug that affected about 4% of predictions; 214,683,829 models
May 8, 2024AlphaFold 3 and AlphaFold Server released
October 9, 2024Nobel Prize in Chemistry awarded in part for AlphaFold
October 2025v6 synchronizes all entries with UniProt 2025_03, adds isoforms and multiple sequence alignments; 241,070,489 models
February 2026AllTheBacteria, Kinetoplastid, Viro3D, and BFVD datasets added
March 20261,754,199 NVIDIA homodimers added
May 2026About 2.2 million homodimers and heterodimers added
September 2026About 1.7 million viral protein complexes, the Viral AlphaFold Database, and Portal View added
Table 2. AlphaFold and AlphaFold DB timeline. Model milestones and database releases. Sources: Google DeepMind (2026), AlphaFold DB FAQ (2026), PDBe release notes (2025), Varadi et al. (2024).[5][1][2][4]

The platform around the models has also grown. The 2025 update added domain annotations from The Encyclopedia of Domains, more than 361 million domains across over 165 million proteins, and reported 4,501,953 website users.[8] Google DeepMind puts AlphaFold's reach at over 3 million researchers in more than 190 countries and the method paper at more than 40,000 citations.[5] Those two user counts measure different things, database visitors and researchers using AlphaFold in any form, and should not be compared directly.

What is not in the AlphaFold Database?

AlphaFold DB models contain only protein atoms. AlphaFold does not predict the positions of ligands, cofactors, metal ions, DNA or RNA, or post-translational modifications, so none of these appear in the downloaded files.[1] Side chains are often arranged as they would be with the missing partner present, for example around a zinc-binding site or a heme pocket, because the model learned from structures in the PDB where those partners were bound. Modelling the partner itself needs a co-folding model such as AlphaFold 3 or open alternatives like Boltz-2, Chai-1, Protenix, and OpenFold-3, which predict proteins together with ligands, ions, and nucleic acids.

Coverage also has hard limits. UniProt sequences are included only if they are between 16 and 2,700 amino acids long for reference proteomes and Swiss-Prot, or up to 1,280 amino acids for the rest of UniProt, a range that includes most proteins, since the median human protein is 415 amino acids long. Only human proteins longer than that are split into overlapping fragments. Sequences with non-standard amino acids, such as X or O, are excluded, as are those added or changed in UniProt after the last synchronization.[1] The broader UniProt predictions come from a single model run, while Swiss-Prot and proteome entries use the most confident of five runs. The difference is small on benchmark targets, about 1 GDT point and a slight tendency toward lower pLDDT.[1]

Each entry is one static conformation of a single chain, unless it comes from a dimer or complex dataset. Where a protein has several known conformations, AlphaFold usually produces only one, and the prediction cannot be steered toward a particular state. Methods such as AF-Cluster, which splits the sequence alignment into subgroups, and AlphaFlow, which samples conformational ensembles, were developed to recover the missing states. AlphaFold has not been validated for predicting the effects of mutations, which dedicated stability predictors such as ThermoMPNN are built for.[1]

Finally, sequences change after they are modelled. Tsitsa and colleagues compared the original human models, built from 2021 sequences, with UniProt 2025_03 and found that 631 of 20,504 human models (3.08%) no longer matched the current sequence. In zebrafish, where curation had revised much of the proteome, the figure was 46.68%.[9] The v6 release resynchronized the core collection with that UniProt version, but the gap will reopen with every later UniProt release.[2] A protein that falls outside the length limits or changed after the last synchronization can be predicted directly with AlphaFold 2 or, faster and without an alignment, ESMFold; our guide to running AlphaFold 2 online walks through the inputs and outputs. Models in the database are licensed under CC-BY 4.0 for academic and commercial use.[1]

How much of the human proteome does AlphaFold cover?

The original AlphaFold human-proteome dataset produced a full-chain prediction for 98.5% of human proteins, but only 58% of residues were predicted with confidence. Before that, after decades of experimental work, only 17% of residues in human protein sequences were covered by an experimentally determined structure.[10]

Those figures use different denominators. Protein-level coverage asks whether a protein received a model at all. Residue-level coverage asks how much of each sequence received a pLDDT above 70. In the 2021 dataset, 36% of all residues reached very high confidence, above 90, and 43.8% of proteins had confident predictions for at least three-quarters of their length.[10] Much of the rest is intrinsically disordered, meaning it does not fold into a fixed shape on its own.

The study used one representative sequence per gene and capped ordinary full-chain predictions at 2,700 amino acids.[10] The database now also includes human isoforms and fragments of longer human proteins, so the 98.5% figure is not a count of every human protein variant. Our human proteome guide explains why gene, protein, isoform, and proteoform counts differ, and the guide to the number of proteins covers the same question across all organisms.

How should AlphaFold confidence scores be interpreted?

AlphaFold reports three kinds of confidence, each answering a different question. pLDDT is a per-residue score from 0 to 100 that measures local confidence: whether a residue sits correctly relative to its neighbours. It does not say whether two domains, or two chains in a complex, are placed correctly relative to each other.[1] For that, AlphaFold DB points to the predicted aligned error (PAE) and, for complexes, to interface scores.

ScoreWhat it measuresRangeInterpretation
pLDDTLocal confidence per residueAbove 90Very high accuracy, suitable for characterizing binding sites
pLDDTLocal confidence per residue70 to 90Confident, generally a good backbone
pLDDTLocal confidence per residue50 to 70Low confidence, treat with caution
pLDDTLocal confidence per residueBelow 50Ribbon-like, often disordered; do not interpret
PAEExpected position error of residue x when aligned on residue y, in ångströmsLow between domainsRelative position and orientation of the domains is well defined
PAEExpected position error of residue x when aligned on residue y, in ångströmsHigh between domainsRelative placement is uncertain and should not be interpreted
ipTMConfidence in the relative positions of chains in a complexAbove 0.8 / 0.6 to 0.8 / below 0.6Confident / uncertain / complex may be incorrect
ipSAEPAE-based interface score over close, confidently placed residue pairs0.8 or more / 0.7 to 0.8 / 0.6 to 0.7 / below 0.6Very high / confident / low / very low
pDockQ2Combines interface pLDDT and PAE into a predicted DockQ scoreAbove about 0.23Likely acceptable interface
Table 3. AlphaFold DB confidence thresholds. Rules of thumb from the official AlphaFold DB documentation; they describe expected reliability, not a guarantee. Source: AlphaFold DB FAQ (2026).[1]

pLDDT is stored in the B-factor column of the PDB and mmCIF files, but unlike a crystallographic B-factor, higher is better. The files contain coordinates for every residue whatever its score, so a low-confidence loop is still present in the file even though it should not be read as real structure.[1] A pLDDT below 50 is a reasonably strong predictor of disorder, meaning a region that is unstructured in isolation or folds only when bound to a partner.[10]

The PAE plot shown on each entry page answers the question pLDDT cannot. Two domains can each have pLDDT above 90 while their relative orientation is essentially a guess, and only high inter-domain PAE reveals it. PAE is also asymmetric: the error at (x, y) can differ from that at (y, x).[1] For complexes, the FAQ recommends reading several interface scores together, because low-pLDDT or disordered regions can pull ipTM down even when the interaction is real. The dimer entries on the website have already been filtered by pDockQ2 and ipSAE, so their scores skew high.[1] The same interface score can be computed for your own predicted complexes with ipSAE.

High-confidence regions support tasks such as structural comparison with FoldSeek or preliminary pocket detection with Fpocket before protein–ligand docking. Experimental evidence is still needed when a conclusion depends on an exact ligand pose, an alternative conformation, a mutation effect, or a flexible interface. AlphaFold 2 returns pLDDT and PAE alongside every model it predicts, so the same thresholds apply to your own runs as to database entries.

Is AlphaFold DB the same as the Protein Data Bank?

No. AlphaFold DB is a database of predictions, while the PDB archive stores structures determined experimentally, by X-ray crystallography, cryo-electron microscopy, or NMR, and deposited by researchers.[11]

RCSB.org shows both kinds of record but keeps them in separately labelled result sets. On September 24, 2026, RCSB listed 259,987 experimental PDB entries and 1,062,058 computed structure models from AlphaFold DB and ModelArchive.[12] Those computed models are a curated subset of about one million AlphaFold predictions for model organisms, global health proteomes, and Swiss-Prot, less than 1% of the full AlphaFold DB.[11] We divided the AlphaFold DB total by the number of PDB entries on the same date, which gives about 1,006 predicted models per experimental entry.

261,552,403 predicted models259,987 PDB entries≈1,006\frac{261{,}552{,}403 \text{ predicted models}}{259{,}987 \text{ PDB entries}} \approx 1{,}006259,987 PDB entries261,552,403 predicted models​≈1,006

The comparison is not like for like. A PDB entry can contain many chains, ligands, and several copies of the same protein, while an AlphaFold DB entry usually describes one chain. The RCSB statistics guide covers the experimental archive and its counting rules.

The scale of AlphaFold DB also made it possible to map protein structure space itself. Clustering all 214 million v4 predictions by structural similarity with Foldseek, the method behind protein structure search, produced 2.30 million clusters with at least two members, and 31% of those clusters had no functional annotation, probably representing previously undescribed structures.[13] Those unannotated clusters are small, covering only 4% of the proteins in the database, so most predicted proteins belong to structural families that already carry some annotation.[13]

Taken together, the numbers describe a resource that is enormous but bounded. It holds a prediction for nearly every catalogued protein sequence, a thousand times more than experiment has solved, yet each model shows one chain in one state without its ligands, and roughly two-fifths of human residues fall below confident prediction. The count of 261.6 million measures how far structural coverage has spread; pLDDT, PAE, and the interface scores measure how much of it can be trusted for any single question.

Sources13
  1. AlphaFold Protein Structure Database: Frequently asked questions

    EMBL-EBI and Google DeepMind · September 24, 2026

  2. AlphaFold Database release notes

    Protein Data Bank in Europe · September 24, 2026

  3. What is the AlphaFold Database?

    EMBL-EBI Training · September 24, 2026

  4. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences

    Nucleic Acids Research · 2024

  5. AlphaFold

    Google DeepMind · September 24, 2026

  6. Highly accurate protein structure prediction with AlphaFold

    Nature · 2021

  7. Accurate structure prediction of biomolecular interactions with AlphaFold 3

    Nature · 2024

  8. AlphaFold Protein Structure Database 2025: a redesigned interface and updated structural coverage

    Nucleic Acids Research · 2025

  9. The aging of the AlphaFold Database

    Nature Structural & Molecular Biology · 2025

  10. Highly accurate protein structure prediction for the human proteome

    Nature · 2021

  11. Computed Structure Models and RCSB.org

    RCSB Protein Data Bank · September 24, 2026

  12. RCSB PDB Search API: entry counts by structure determination methodology

    RCSB Protein Data Bank · September 24, 2026

  13. Clustering predicted structures at the scale of the known protein universe

    Nature · 2023

Cite this article

Broz, M. (2026, September 24). AlphaFold database statistics [2026]. ProteinIQ. https://proteiniq.io/guides/alphafold-statistics

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “AlphaFold database statistics [2026]” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
July 1, 2026
Updated
September 24, 2026

Related guides

Browse all guides
Ink illustration of archived protein structure records and a crystal representing the Protein Data Bank.

Structures · August 9, 2026

Protein Data Bank statistics 2026: size and coverage

Current Protein Data Bank statistics for archive size, protein sequence coverage, experimental methods, growth, usage, computed models, and structure data formats.

Two protein chains with low-PAE residue pairs highlighted at their interface and faint flexible tails

Structures · October 2, 2026

ipSAE explained

ipSAE is an interface confidence score for AlphaFold2, AlphaFold3 and Boltz predictions that counts only confidently placed residue pairs between chains. Learn how it is calculated, why it beats ipTM on full-length sequences, what 0.6, 0.7 and 0.8 mean, and how to calculate it for your own complexes and binder designs.

Two-chain protein complex with pTM marking the overall structure and ipTM marking relationships between chains

Structures · October 1, 2026

ipTM vs pTM: judging predicted complexes

pTM scores the whole predicted structure and ipTM scores only the placement of chains relative to each other. Learn how both are calculated, what the 0.5, 0.6 and 0.8 cutoffs mean, why disordered tails lower ipTM, and how to read them in AlphaFold2, Boltz-2 and other complex predictors.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing