# Proteome sizes by species: protein counts compared

> Compare protein-coding genes, UniProtKB entries, and separate isoforms across 12 reference proteomes, with scope-aware examples of large and small proteomes.

Proteome size spans nearly three orders of magnitude across the organisms in this comparison. The current one-protein-per-gene count runs from 137 in a host-dependent bacterial endosymbiont to 105,080 in bread wheat.

The ranking changes when “protein count” means database entries rather than genes. This guide keeps protein-coding genes, UniProtKB entries, and separately stored isoforms in different columns.

## How large are proteomes across species?

Across these 12 UniProt reference proteomes, the one-per-gene count ranges from 137 protein sequences in _Candidatus Nasuia deltocephalinicola_ to 105,080 in bread wheat.

| Organism                                       | Scope                        | Protein-coding genes | UniProtKB entries | Separate isoforms |
| ---------------------------------------------- | ---------------------------- | -------------------- | ----------------- | ----------------- |
| Bread wheat                                    | Hexaploid crop               | 105,080              | 130,692           | 0                 |
| _Arabidopsis thaliana_                         | Diploid plant model          | 27,496               | 39,273            | 2,324             |
| Zebrafish                                      | Vertebrate model             | 26,677               | 68,855            | 268               |
| Mouse                                          | Mammal model                 | 21,853               | 54,857            | 8,467             |
| Human                                          | Human reference              | 20,652               | 147,506           | 22,131            |
| _Caenorhabditis elegans_                       | Nematode model               | 19,789               | 26,629            | 1,802             |
| Fruit fly                                      | Insect model                 | 13,817               | 21,953            | 1,533             |
| Budding yeast                                  | Unicellular eukaryote        | 6,066                | 6,067             | 31                |
| _Escherichia coli_ K-12                        | Bacterial model strain       | 4,403                | 4,403             | 12                |
| _Pelagibacter ubique_ HTCC1062                 | Free-living marine bacterium | 1,354                | 1,354             | 0                 |
| _Mycoplasmoides genitalium_ G37                | Obligate parasite            | 483                  | 483               | 1                 |
| _Candidatus Nasuia deltocephalinicola_ NAS-ALF | Obligate endosymbiont        | 137                  | 137               | 0                 |

All rows come from UniProt release 2026_02 and use one named reference proteome per organism or strain. UniProt defines reference proteomes as representative landmarks selected to cover the tree of life. They do not cover every proteome or species.

![Protein-coding gene counts for seven selected UniProt reference proteomes, led by hexaploid bread wheat](/images/charts/protein-coding-genes-selected-reference-proteomes.webp)

These counts measure how many coding genes are represented in each selected genome annotation. Protein molecule abundance, tissue expression, and modified [proteoforms](/guides/human-proteome-size) require different measurements.

## Which species has the largest proteome?

There is no defensible single “largest proteome” without defining the organisms searched and the counting unit. In this selected set, bread wheat has the most protein-coding genes at 105,080, while the human reference proteome has the most UniProtKB entries at 147,506.

Bread wheat is hexaploid: its genome combines A, B, and D subgenomes. The 2018 reference assembly originally reported 107,891 high-confidence gene models, while the current UniProt reference record lists 105,080 protein-coding genes under its present annotation. That difference shows why a proteome statistic needs a release and assembly, even when the species name stays the same.

Human moves ahead only when the comparison switches to UniProtKB entries. The 147,506 figure counts entries rather than distinct protein-coding genes. For the narrower human question, the [human proteome size](/guides/human-proteome-size) guide compares current gene catalogs, reference proteins, translations, and proteoforms.

## Which organism has the smallest proteome?

The smallest proteome depends on whether obligate endosymbionts, parasites, and free-living organisms are compared together. Within this dataset, the smallest counts are 137 genes for an obligate endosymbiont, 483 for an obligate parasite, and 1,354 for a free-living marine bacterium.

![Protein-coding gene counts in E. coli and three organisms used in small-proteome comparisons, separated by lifestyle](/images/charts/reduced-reference-proteomes-by-lifestyle.webp)

_Candidatus Nasuia deltocephalinicola_ lives inside a leafhopper and depends on its host and a second bacterial symbiont. Its 112-kilobase genome was reported as the smallest bacterial genome sequenced at the time; the current UniProt reference record contains 137 protein-coding genes. Its dependence on other organisms makes it an endosymbiont extreme and rules it out as a self-sufficient cell benchmark.

_Mycoplasmoides genitalium_ is an obligate human parasite that can be grown in pure culture. A genome-wide essentiality study described 482 protein-coding genes, and the current reference proteome lists 483.

_Pelagibacter ubique_ HTCC1062 provides a free-living comparison. Its original genome study reported 1,354 predicted open reading frames and complete pathways for all 20 standard amino acids; UniProt still lists 1,354 protein-coding genes for this reference proteome.

## Why can protein counts exceed gene counts?

Protein-entry counts can exceed gene counts because one gene may be represented by several protein sequences, and some alternative isoforms are stored separately from the main UniProtKB entry.

UniProtKB/Swiss-Prot chooses one canonical sequence for display in each curated entry and describes alternative products with that entry when possible. UniProtKB/TrEMBL can also contain additional predicted sequences for genes already represented by a curated entry. Adding “UniProtKB entries” and “separate isoforms” would produce a sequence-record total, not a count of genes or distinct protein products.

The human row makes the distinction visible: 20,652 genes, 147,506 UniProtKB entries, and 22,131 separately counted isoform sequences in this release. Other resources use different annotation rules. The [number of proteins](/guides/number-of-proteins) guide explains how gene, sequence, structure, proteoform, and molecule counts answer different questions.

## How were these proteome sizes compared?

The comparison uses the `geneCount`, `proteinCount`, and `isoformProteinCount` fields returned for 12 named UniProt reference proteomes in release 2026_02, accessed on August 9, 2026.

The charts use `geneCount` because UniProt uses that field as the basis for downloading one protein sequence per gene. This gives the closest common baseline across bacteria and eukaryotes. The table preserves all three database fields so readers can see where annotation expands the protein-entry count.

[Download the organism-level CSV](/data/guides/proteome-sizes-by-species-2026-08-09.csv) for the exact values, proteome identifiers, assemblies, annotation sources, record dates, and source URLs used here. Protein length is a separate property; the [average protein size](/guides/average-protein-size) comparison uses sequence length rather than the number of coding genes. The [E. coli statistics](/guides/e-coli-statistics) guide covers the K-12 strain's gene and protein counts in more detail.
