ProteinIQ
Sign inStart for free
ProteinIQ
Statistics/Aug '26/3 min read

AlphaFold database statistics [2026]

Matic Broz

Matic BrozComputational chemist

The AlphaFold Protein Structure Database contains 262,739,159 predicted models as of August 9, 2026. That total includes the main UniProt collection as well as isoforms, fragments, protein complexes, and community datasets.

This is a database count, not a count of unique protein types or experimentally determined structures.

How many structures are in the AlphaFold Database?

AlphaFold DB currently lists 262,739,159 predicted models, including 40,054 isoforms, and provides 46 complete proteomes for bulk download.[1]

The latest numbered core release is v6, which contains 241,070,489 predictions synchronized to UniProt 2025_03.[3] The larger website total also includes collections added after that release. These include more than 17.7 million AllTheBacteria predictions, viral datasets, fragments of long proteins, and 2,158,419 homodimer plus 79,156 heterodimer models from the NVIDIA collection.[1] One UniProt accession can therefore have several database entries.

A separate EMBL-EBI training page still listed 260,986,406 structures on August 9, 2026.[2] The live AlphaFold DB FAQ showed 262,739,159 on the same date, so this article uses the database-facing figure and attaches an access date.[1]

AlphaFold Database growth from more than 360,000 models at launch to 262.7 million models on the current website

AlphaFold DB launched in July 2021 with more than 360,000 structures for 20 model-organism proteomes. Version 4 contained 214,683,829 predictions, and version 6 raised the core UniProt collection to 241,070,489.[1][3][4] Researchers can retrieve the current model files through the AlphaFold database downloader.

How much of the human proteome does AlphaFold cover?

The original AlphaFold human-proteome dataset produced a full-chain prediction for 98.5% of human proteins, but only 58.0% of residues were predicted with confident local structure.[5]

Those figures describe different denominators. Protein-level coverage asks whether a protein received a model. Residue-level coverage asks how much of each sequence received pLDDT above 70. In the 2021 dataset, 35.7% of all residues were in the highest-confidence band above 90, and 43.8% of proteins had confident predictions for at least three-quarters of their sequence.[5]

The study used one representative sequence per gene and capped ordinary full-chain predictions at 2,700 amino acids.[5] The current database also includes human isoforms and segmented models for longer human proteins, so the published 98.5% result should not be treated as a count of every human proteoform. Our human proteome guide explains why gene, protein, isoform, and proteoform counts differ.

How should AlphaFold confidence scores be interpreted?

pLDDT above 90 usually marks a locally high-accuracy region, while scores above 70 generally indicate a reliable backbone. Scores from 50 to 70 indicate low confidence, and regions below 50 often correspond to disorder or structure that depends on a binding partner.[1][5]

pLDDT is a residue-level confidence score, not a probability that the protein adopts that structure in every biological state. AlphaFold DB supplies coordinates for low-confidence regions too, so their presence in a PDB or mmCIF file does not make them reliable. Predicted aligned error, or PAE, is the better measure for judging the relative placement of domains.

High-confidence regions can support tasks such as structural comparison with FoldSeek or preliminary pocket analysis with Fpocket. Experimental evidence is still needed when a conclusion depends on an exact ligand pose, alternative conformation, mutation effect, or flexible interface. ProteinIQ's AlphaFold 2 implementation exposes the model and confidence outputs together.

Is AlphaFold DB the same as the Protein Data Bank?

No. AlphaFold DB is a prediction database, while the PDB archive stores experimentally determined and integrative structures submitted by researchers.[6]

RCSB.org displays both kinds of records, but it keeps them in separate result sets. On August 9, 2026, RCSB listed 258,023 PDB archive structures and 1,062,058 computed structure models from AlphaFold DB and ModelArchive.[6] RCSB documents that only about one million AlphaFold models are integrated into its site, a small subset of the 262.7 million models available from the full AlphaFold protein structure database.[1][6]

The same protein can have multiple experimental entries, predicted models, conformations, ligands, and sequence variants. The RCSB statistics guide covers the experimental archive and its counting rules in detail.

Sources▼
  1. AlphaFold Protein Structure Database: Frequently asked questions EMBL-EBI and Google DeepMind · August 9, 2026. https://alphafold.ebi.ac.uk/faq
  2. What is the AlphaFold Database? EMBL-EBI Training · August 9, 2026. https://www.ebi.ac.uk/training/online/courses/navigating-alphafold-database/what-is-the-afdb/
  3. AlphaFold Database release notes Protein Data Bank in Europe · August 9, 2026. https://www.ebi.ac.uk/pdbe/news/alphafold-database-release-notes
  4. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences Nucleic Acids Research · 2024. https://academic.oup.com/nar/article/52/D1/D368/7337620
  5. Highly accurate protein structure prediction for the human proteome Nature · 2021. https://www.nature.com/articles/s41586-021-03828-1
  6. Computed Structure Models and RCSB.org RCSB Protein Data Bank · August 9, 2026. https://www.rcsb.org/docs/general-help/computed-structure-models-and-rcsborg
Published
July 1, 2026
Last updated
August 9, 2026

Table of contents

Cite this article

Broz, M. (2026, August 9). AlphaFold database statistics [2026]. ProteinIQ. https://proteiniq.io/guides/alphafold-statistics

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

Related guides

Statistics

AUG '26

Protein Data Bank statistics 2026: size and coverage

Current Protein Data Bank statistics for archive size, protein sequence coverage, experimental methods, growth, usage, computed models, and structure data formats.

Matic Broz Computational chemist

Statistics

JUL '26

AI drug discovery statistics [2026]

AI drug discovery statistics in 2026 include more than 500 FDA submissions with AI components from 2016 to 2023, 75 AI-discovered molecules entered the clinic by 2023, 241 million AlphaFold DB predictions, and measurable virtual-screening and ADMET benchmarks.

Matic Broz Computational chemist

Statistics

AUG '26

How many types of antibodies are there?

The human body has five antibody classes. Compare the structures, locations, functions, and subclasses of IgG, IgA, IgM, IgD, and IgE.

Matic Broz Computational chemist

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing