# AlphaFold database statistics [2026]

> AlphaFold DB lists 261,552,403 predicted models as of September 24, 2026. See the collection breakdown, release timeline, human proteome coverage, pLDDT, PAE and interface thresholds, and how the database compares with the PDB.

The AlphaFold Protein Structure Database lists **261,552,403 predicted protein models** as of September 24, 2026. About 241 million of them form the core release, one [AlphaFold 2](/app/alphafold-2) prediction for most sequences in UniProt, the central catalogue of [known proteins](/guides/number-of-proteins). The rest come from partner datasets that add bacterial genomes, viruses, parasites, and predicted protein pairs.

That figure counts computer predictions, not experimentally determined structures, and one protein can have several entries. For scale, the [Protein Data Bank](/guides/rcsb-statistics) holds about 260,000 experimental entries, so AlphaFold DB offers roughly a thousand predictions for every structure solved in a laboratory.

## How many structures are in the AlphaFold Database?

The AlphaFold DB FAQ lists 261,552,403 predicted models, including 40,054 isoforms, with 46 complete [proteomes](/guides/proteome-sizes-by-species) available for bulk download. The latest numbered core release, v6 from October 2025, contains 241,070,489 predictions synchronized to UniProt release 2025_03. The difference between the two numbers is the partner collections added since then.

  <caption><strong>Table 1. What the AlphaFold DB total is made of.</strong> Collection sizes as reported by AlphaFold DB on September 24, 2026. For the dimer datasets, the counts are the models shown on entry pages after filtering by interface confidence (pDockQ2 of at least 0.23 and ipSAE of at least 0.6). Sources: AlphaFold DB FAQ (2026), PDBe release notes (2025).</caption>
  <thead>
    <tr>
      <th scope="col">Collection</th>
      <th scope="col">Predicted models</th>
      <th scope="col">What it covers</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>UniProt core release (v6)</td>
      <td>241,070,489</td>
      <td>Single chains for UniProt 2025_03, including 40,054 isoforms</td>
    </tr>
    <tr>
      <td>AllTheBacteria</td>
      <td>over 17.7 million</td>
      <td>Hypothetical proteins from more than 2.4 million bacterial and archaeal genomes</td>
    </tr>
    <tr>
      <td>NVIDIA proteome-scale dimers</td>
      <td>2,158,419 homodimers and 79,156 heterodimers</td>
      <td>Protein pairs from Swiss-Prot, 16 model organisms, and 30 WHO global health proteomes</td>
    </tr>
    <tr>
      <td>Viral complexes</td>
      <td>1,703,992</td>
      <td>Pairwise viral protein interactions across about 2,800 viral proteomes</td>
    </tr>
    <tr>
      <td>Big Fantastic Virus Database (BFVD)</td>
      <td>351,242</td>
      <td>Viral protein sequences from UniRef30 2023_02</td>
    </tr>
    <tr>
      <td>Viro3D</td>
      <td>82,424</td>
      <td>Proteins from 4,407 human and animal viruses</td>
    </tr>
    <tr>
      <td>Viral AlphaFold Database (VAD)</td>
      <td>53,716</td>
      <td>Full-length viral RefSeq proteins, as monomers and homodimers</td>
    </tr>
    <tr>
      <td>Kinetoplastids</td>
      <td>16,887</td>
      <td>Trypanosoma and Leishmania proteins</td>
    </tr>
  </tbody>

By our tally, the collection sizes excluding the viral complexes add up to about 261.5 million, within 0.02% of the headline figure. The viral-complex release arrived in September 2026 and appears not yet to be included in the FAQ counter, although the FAQ does not say which collections its total covers. The count is also revised downward at times. In August 2026 the same FAQ listed 262,739,159 models, and an EMBL-EBI training page still showed 260,986,406 on September 24, 2026. Anyone citing a single number should give the page and the access date.

These totals count database entries, not distinct proteins. One UniProt accession can have a canonical model, isoform models (a hyphenated suffix such as P42167-2), fragments of a long protein (such as Q8WZ42-F1, a piece of titin, the [largest human protein](/guides/largest-protein)), and separate predictions from partner datasets made with different software.

![AlphaFold Database growth from more than 360,000 models at launch to 261.6 million models in September 2026](/images/charts/alphafold-db-release-growth.webp '**Figure 1. Growth of the AlphaFold Database.** Points mark the launch, the v4 and v6 core releases, and the website total on September 24, 2026, and are not evenly spaced in time. Sources: [Varadi et al. (2024)](https://academic.oup.com/nar/article/52/D1/D368/7337620), [PDBe release notes](https://www.ebi.ac.uk/pdbe/news/alphafold-database-release-notes), [AlphaFold DB FAQ](https://alphafold.ebi.ac.uk/faq).')

AlphaFold DB launched in July 2021 with more than 360,000 structures from 20 model-organism proteomes, and version 4 contained 214,683,829 predictions. Nearly all of that growth happened in a single step in July 2022, when most of UniProt was added at once. Since then the core has grown slowly with UniProt, and most new entries come from partner datasets. The full data can be retrieved from the EBI FTP site and Google Cloud; the v4 Google Cloud dataset alone was about 23 TiB. Individual model files can also be fetched with the [AlphaFold database downloader](/app/alphafold-database-download).

## When was AlphaFold released, and how has the database changed?

AlphaFold 2 was presented at the CASP14 prediction assessment on November 30, 2020, and its method paper and source code were published on July 15, 2021. The database followed a week later, on July 22, 2021. The database still contains AlphaFold 2 and related ColabFold predictions. AlphaFold 3, released with the AlphaFold Server on May 8, 2024, is not the source of its models.

  <caption><strong>Table 2. AlphaFold and AlphaFold DB timeline.</strong> Model milestones and database releases. Sources: Google DeepMind (2026), AlphaFold DB FAQ (2026), PDBe release notes (2025), Varadi et al. (2024).</caption>
  <thead>
    <tr>
      <th scope="col">Date</th>
      <th scope="col">Event</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>November 30, 2020</td>
      <td>AlphaFold 2 tops the CASP14 assessment</td>
    </tr>
    <tr>
      <td>July 15, 2021</td>
      <td>AlphaFold 2 paper published in Nature and code released</td>
    </tr>
    <tr>
      <td>July 22, 2021</td>
      <td>AlphaFold DB launches with 20 model-organism proteomes, more than 360,000 structures</td>
    </tr>
    <tr>
      <td>December 2021</td>
      <td>Swiss-Prot added</td>
    </tr>
    <tr>
      <td>January 2022</td>
      <td>WHO global health proteomes added</td>
    </tr>
    <tr>
      <td>July 28, 2022</td>
      <td>Most of UniProt added, taking the database past 200 million structures</td>
    </tr>
    <tr>
      <td>November 2022</td>
      <td>v4 fixes a numerical bug that affected about 4% of predictions; 214,683,829 models</td>
    </tr>
    <tr>
      <td>May 8, 2024</td>
      <td>AlphaFold 3 and AlphaFold Server released</td>
    </tr>
    <tr>
      <td>October 9, 2024</td>
      <td>Nobel Prize in Chemistry awarded in part for AlphaFold</td>
    </tr>
    <tr>
      <td>October 2025</td>
      <td>v6 synchronizes all entries with UniProt 2025_03, adds isoforms and multiple sequence alignments; 241,070,489 models</td>
    </tr>
    <tr>
      <td>February 2026</td>
      <td>AllTheBacteria, Kinetoplastid, Viro3D, and BFVD datasets added</td>
    </tr>
    <tr>
      <td>March 2026</td>
      <td>1,754,199 NVIDIA homodimers added</td>
    </tr>
    <tr>
      <td>May 2026</td>
      <td>About 2.2 million homodimers and heterodimers added</td>
    </tr>
    <tr>
      <td>September 2026</td>
      <td>About 1.7 million viral protein complexes, the Viral AlphaFold Database, and Portal View added</td>
    </tr>
  </tbody>

The platform around the models has also grown. The 2025 update added domain annotations from The Encyclopedia of Domains, more than 361 million domains across over 165 million proteins, and reported 4,501,953 website users. Google DeepMind puts AlphaFold's reach at over 3 million researchers in more than 190 countries and the method paper at more than 40,000 citations. Those two user counts measure different things, database visitors and researchers using AlphaFold in any form, and should not be compared directly.

## What is not in the AlphaFold Database?

AlphaFold DB models contain only protein atoms. AlphaFold does not predict the positions of ligands, cofactors, metal ions, DNA or RNA, or post-translational modifications, so none of these appear in the downloaded files. Side chains are often arranged as they would be with the missing partner present, for example around a zinc-binding site or a heme pocket, because the model learned from structures in the PDB where those partners were bound. Modelling the partner itself needs a co-folding model such as AlphaFold 3 or open alternatives like [Boltz-2](/app/boltz-2), [Chai-1](/app/chai-1), [Protenix](/app/protenix), and [OpenFold-3](/app/openfold-3), which predict proteins together with ligands, ions, and nucleic acids.

Coverage also has hard limits. UniProt sequences are included only if they are between 16 and 2,700 amino acids long for reference proteomes and Swiss-Prot, or up to 1,280 amino acids for the rest of UniProt, a range that includes most proteins, since the [median human protein](/guides/average-protein-size) is 415 amino acids long. Only human proteins longer than that are split into overlapping fragments. Sequences with [non-standard amino acids](/guides/how-many-amino-acids-are-there), such as X or O, are excluded, as are those added or changed in UniProt after the last synchronization. The broader UniProt predictions come from a single model run, while Swiss-Prot and proteome entries use the most confident of five runs. The difference is small on benchmark targets, about 1 GDT point and a slight tendency toward lower pLDDT.

Each entry is one static conformation of a single chain, unless it comes from a dimer or complex dataset. Where a protein has several known conformations, AlphaFold usually produces only one, and the prediction cannot be steered toward a particular state. Methods such as [AF-Cluster](/app/af-cluster), which splits the sequence alignment into subgroups, and [AlphaFlow](/app/alphaflow), which samples conformational ensembles, were developed to recover the missing states. AlphaFold has not been validated for predicting the effects of mutations, which dedicated stability predictors such as [ThermoMPNN](/app/thermompnn) are built for.

Finally, sequences change after they are modelled. Tsitsa and colleagues compared the original human models, built from 2021 sequences, with UniProt 2025_03 and found that 631 of 20,504 human models (3.08%) no longer matched the current sequence. In zebrafish, where curation had revised much of the proteome, the figure was 46.68%. The v6 release resynchronized the core collection with that UniProt version, but the gap will reopen with every later UniProt release. A protein that falls outside the length limits or changed after the last synchronization can be predicted directly with [AlphaFold 2](/app/alphafold-2) or, faster and without an alignment, [ESMFold](/app/esmfold); our guide to [running AlphaFold 2 online](/guides/how-to-use-alphafold2-online) walks through the inputs and outputs. Models in the database are licensed under CC-BY 4.0 for academic and commercial use.

## How much of the human proteome does AlphaFold cover?

The original AlphaFold human-proteome dataset produced a full-chain prediction for 98.5% of human proteins, but only 58% of residues were predicted with confidence. Before that, after decades of experimental work, only 17% of residues in human protein sequences were covered by an experimentally determined structure.

Those figures use different denominators. Protein-level coverage asks whether a protein received a model at all. Residue-level coverage asks how much of each sequence received a pLDDT above 70. In the 2021 dataset, 36% of all residues reached very high confidence, above 90, and 43.8% of proteins had confident predictions for at least three-quarters of their length. Much of the rest is intrinsically disordered, meaning it does not fold into a fixed shape on its own.

The study used one representative sequence per gene and capped ordinary full-chain predictions at 2,700 amino acids. The database now also includes human isoforms and fragments of longer human proteins, so the 98.5% figure is not a count of every human protein variant. Our [human proteome](/guides/human-proteome-size) guide explains why gene, protein, isoform, and proteoform counts differ, and the guide to the [number of proteins](/guides/number-of-proteins) covers the same question across all organisms.

## How should AlphaFold confidence scores be interpreted?

AlphaFold reports three kinds of confidence, each answering a different question. pLDDT is a per-residue score from 0 to 100 that measures local confidence: whether a residue sits correctly relative to its neighbours. It does not say whether two domains, or two chains in a complex, are placed correctly relative to each other. For that, AlphaFold DB points to the predicted aligned error (PAE) and, for complexes, to interface scores.

  <caption><strong>Table 3. AlphaFold DB confidence thresholds.</strong> Rules of thumb from the official AlphaFold DB documentation; they describe expected reliability, not a guarantee. Source: AlphaFold DB FAQ (2026).</caption>
  <thead>
    <tr>
      <th scope="col">Score</th>
      <th scope="col">What it measures</th>
      <th scope="col">Range</th>
      <th scope="col">Interpretation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>pLDDT</td>
      <td>Local confidence per residue</td>
      <td>Above 90</td>
      <td>Very high accuracy, suitable for characterizing binding sites</td>
    </tr>
    <tr>
      <td>pLDDT</td>
      <td>Local confidence per residue</td>
      <td>70 to 90</td>
      <td>Confident, generally a good backbone</td>
    </tr>
    <tr>
      <td>pLDDT</td>
      <td>Local confidence per residue</td>
      <td>50 to 70</td>
      <td>Low confidence, treat with caution</td>
    </tr>
    <tr>
      <td>pLDDT</td>
      <td>Local confidence per residue</td>
      <td>Below 50</td>
      <td>Ribbon-like, often disordered; do not interpret</td>
    </tr>
    <tr>
      <td>PAE</td>
      <td>Expected position error of residue x when aligned on residue y, in ångströms</td>
      <td>Low between domains</td>
      <td>Relative position and orientation of the domains is well defined</td>
    </tr>
    <tr>
      <td>PAE</td>
      <td>Expected position error of residue x when aligned on residue y, in ångströms</td>
      <td>High between domains</td>
      <td>Relative placement is uncertain and should not be interpreted</td>
    </tr>
    <tr>
      <td>ipTM</td>
      <td>Confidence in the relative positions of chains in a complex</td>
      <td>Above 0.8 / 0.6 to 0.8 / below 0.6</td>
      <td>Confident / uncertain / complex may be incorrect</td>
    </tr>
    <tr>
      <td>ipSAE</td>
      <td>PAE-based interface score over close, confidently placed residue pairs</td>
      <td>0.8 or more / 0.7 to 0.8 / 0.6 to 0.7 / below 0.6</td>
      <td>Very high / confident / low / very low</td>
    </tr>
    <tr>
      <td>pDockQ2</td>
      <td>Combines interface pLDDT and PAE into a predicted DockQ score</td>
      <td>Above about 0.23</td>
      <td>Likely acceptable interface</td>
    </tr>
  </tbody>

pLDDT is stored in the B-factor column of the PDB and mmCIF files, but unlike a crystallographic B-factor, higher is better. The files contain coordinates for every residue whatever its score, so a low-confidence loop is still present in the file even though it should not be read as real structure. A pLDDT below 50 is a reasonably strong predictor of disorder, meaning a region that is unstructured in isolation or folds only when bound to a partner.

The PAE plot shown on each entry page answers the question pLDDT cannot. Two domains can each have pLDDT above 90 while their relative orientation is essentially a guess, and only high inter-domain PAE reveals it. PAE is also asymmetric: the error at (x, y) can differ from that at (y, x). For [complexes](/use-cases/protein-complex-structure-prediction), the FAQ recommends reading several interface scores together, because low-pLDDT or disordered regions can pull ipTM down even when the interaction is real. The dimer entries on the website have already been filtered by pDockQ2 and ipSAE, so their scores skew high. The same interface score can be computed for your own predicted complexes with [ipSAE](/app/ipsae).

High-confidence regions support tasks such as structural comparison with [FoldSeek](/app/foldseek) or preliminary pocket detection with [Fpocket](/app/fpocket) before [protein–ligand docking](/use-cases/protein-ligand-docking). Experimental evidence is still needed when a conclusion depends on an exact ligand pose, an alternative conformation, a mutation effect, or a flexible interface. [AlphaFold 2](/app/alphafold-2) returns pLDDT and PAE alongside every model it predicts, so the same thresholds apply to your own runs as to database entries.

## Is AlphaFold DB the same as the Protein Data Bank?

No. AlphaFold DB is a database of predictions, while the PDB archive stores structures determined experimentally, by X-ray crystallography, cryo-electron microscopy, or NMR, and deposited by researchers.

RCSB.org shows both kinds of record but keeps them in separately labelled result sets. On September 24, 2026, RCSB listed 259,987 experimental PDB entries and 1,062,058 computed structure models from AlphaFold DB and ModelArchive. Those computed models are a curated subset of about one million AlphaFold predictions for model organisms, global health proteomes, and Swiss-Prot, less than 1% of the full AlphaFold DB. We divided the AlphaFold DB total by the number of PDB entries on the same date, which gives about 1,006 predicted models per experimental entry.

$$
\frac{261{,}552{,}403 \text{ predicted models}}{259{,}987 \text{ PDB entries}}
\approx 1{,}006
$$

The comparison is not like for like. A PDB entry can contain many chains, ligands, and several copies of the same protein, while an AlphaFold DB entry usually describes one chain. The [RCSB statistics](/guides/rcsb-statistics) guide covers the experimental archive and its counting rules.

The scale of AlphaFold DB also made it possible to map protein structure space itself. Clustering all 214 million v4 predictions by structural similarity with Foldseek, the method behind [protein structure search](/use-cases/protein-structure-search), produced 2.30 million clusters with at least two members, and 31% of those clusters had no functional annotation, probably representing previously undescribed structures. Those unannotated clusters are small, covering only 4% of the proteins in the database, so most predicted proteins belong to structural families that already carry some annotation.

Taken together, the numbers describe a resource that is enormous but bounded. It holds a prediction for nearly every catalogued protein sequence, a thousand times more than experiment has solved, yet each model shows one chain in one state without its ligands, and roughly two-fifths of human residues fall below confident prediction. The count of 261.6 million measures how far structural coverage has spread; pLDDT, PAE, and the interface scores measure how much of it can be trusted for any single question.
