Proteome sizes by species: protein counts compared

Matic BrozComputational chemist
Proteome size spans nearly three orders of magnitude across the organisms in this comparison. The current one-protein-per-gene count runs from 137 in a host-dependent bacterial endosymbiont to 105,080 in bread wheat.
The ranking changes when “protein count” means database entries rather than genes. This guide keeps protein-coding genes, UniProtKB entries, and separately stored isoforms in different columns.
How large are proteomes across species?
Across these 12 UniProt reference proteomes, the one-per-gene count ranges from 137 protein sequences in Candidatus Nasuia deltocephalinicola to 105,080 in bread wheat.[1]
| Organism | Scope | Protein-coding genes | UniProtKB entries | Separate isoforms |
|---|---|---|---|---|
| Bread wheat | Hexaploid crop | 105,080 | 130,692 | 0 |
| Arabidopsis thaliana | Diploid plant model | 27,496 | 39,273 | 2,324 |
| Zebrafish | Vertebrate model | 26,677 | 68,855 | 268 |
| Mouse | Mammal model | 21,853 | 54,857 | 8,467 |
| Human | Human reference | 20,652 | 147,506 | 22,131 |
| Caenorhabditis elegans | Nematode model | 19,789 | 26,629 | 1,802 |
| Fruit fly | Insect model | 13,817 | 21,953 | 1,533 |
| Budding yeast | Unicellular eukaryote | 6,066 | 6,067 | 31 |
| Escherichia coli K-12 | Bacterial model strain | 4,403 | 4,403 | 12 |
| Pelagibacter ubique HTCC1062 | Free-living marine bacterium | 1,354 | 1,354 | 0 |
| Mycoplasmoides genitalium G37 | Obligate parasite | 483 | 483 | 1 |
| Candidatus Nasuia deltocephalinicola NAS-ALF | Obligate endosymbiont | 137 | 137 | 0 |
All rows come from UniProt release 2026_02 and use one named reference proteome per organism or strain.[1] UniProt defines reference proteomes as representative landmarks selected to cover the tree of life. They do not cover every proteome or species.[2]
These counts measure how many coding genes are represented in each selected genome annotation. Protein molecule abundance, tissue expression, and modified proteoforms require different measurements.
Which species has the largest proteome?
There is no defensible single “largest proteome” without defining the organisms searched and the counting unit. In this selected set, bread wheat has the most protein-coding genes at 105,080, while the human reference proteome has the most UniProtKB entries at 147,506.[1]
Bread wheat is hexaploid: its genome combines A, B, and D subgenomes. The 2018 reference assembly originally reported 107,891 high-confidence gene models, while the current UniProt reference record lists 105,080 protein-coding genes under its present annotation.[4][1] That difference shows why a proteome statistic needs a release and assembly, even when the species name stays the same.
Human moves ahead only when the comparison switches to UniProtKB entries. The 147,506 figure counts entries rather than distinct protein-coding genes. For the narrower human question, the human proteome size guide compares current gene catalogs, reference proteins, translations, and proteoforms.
Which organism has the smallest proteome?
The smallest proteome depends on whether obligate endosymbionts, parasites, and free-living organisms are compared together. Within this dataset, the smallest counts are 137 genes for an obligate endosymbiont, 483 for an obligate parasite, and 1,354 for a free-living marine bacterium.[1]
Candidatus Nasuia deltocephalinicola lives inside a leafhopper and depends on its host and a second bacterial symbiont. Its 112-kilobase genome was reported as the smallest bacterial genome sequenced at the time; the current UniProt reference record contains 137 protein-coding genes.[7][1] Its dependence on other organisms makes it an endosymbiont extreme and rules it out as a self-sufficient cell benchmark.
Mycoplasmoides genitalium is an obligate human parasite that can be grown in pure culture. A genome-wide essentiality study described 482 protein-coding genes, and the current reference proteome lists 483.[6][1]
Pelagibacter ubique HTCC1062 provides a free-living comparison. Its original genome study reported 1,354 predicted open reading frames and complete pathways for all 20 standard amino acids; UniProt still lists 1,354 protein-coding genes for this reference proteome.[5][1]
Why can protein counts exceed gene counts?
Protein-entry counts can exceed gene counts because one gene may be represented by several protein sequences, and some alternative isoforms are stored separately from the main UniProtKB entry.[1][3]
UniProtKB/Swiss-Prot chooses one canonical sequence for display in each curated entry and describes alternative products with that entry when possible. UniProtKB/TrEMBL can also contain additional predicted sequences for genes already represented by a curated entry.[3] Adding “UniProtKB entries” and “separate isoforms” would produce a sequence-record total, not a count of genes or distinct protein products.
The human row makes the distinction visible: 20,652 genes, 147,506 UniProtKB entries, and 22,131 separately counted isoform sequences in this release.[1] Other resources use different annotation rules. The number of proteins guide explains how gene, sequence, structure, proteoform, and molecule counts answer different questions.
How were these proteome sizes compared?
The comparison uses the geneCount, proteinCount, and isoformProteinCount fields returned for 12 named UniProt reference proteomes in release 2026_02, accessed on August 9, 2026.[1]
The charts use geneCount because UniProt exposes that field as the basis for downloading one protein sequence per gene. This gives the closest common baseline across bacteria and eukaryotes. The table preserves all three database fields so readers can see where annotation expands the protein-entry count.
Download the organism-level CSV for the exact values, proteome identifiers, assemblies, annotation sources, record dates, and source URLs used here. Protein length is a separate property; the average protein size comparison uses sequence length rather than the number of coding genes. The E. coli statistics guide covers the K-12 strain's gene and protein counts in more detail.
Sources▼
- Selected reference proteomes, UniProt release 2026_02 UniProt · August 9, 2026. https://rest.uniprot.org/proteomes/search?query=%28upid%3AUP000019116%20OR%20upid%3AUP000005640%20OR%20upid%3AUP000000589%20OR%20upid%3AUP000000437%20OR%20upid%3AUP000006548%20OR%20upid%3AUP000001940%20OR%20upid%3AUP000000803%20OR%20upid%3AUP000002311%20OR%20upid%3AUP000000625%20OR%20upid%3AUP000000807%20OR%20upid%3AUP000002528%20OR%20upid%3AUP000015382%29&format=json&size=12
- What are reference proteomes? UniProt · August 9, 2026. https://www.uniprot.org/help/reference_proteome
- What is the canonical sequence? Are all isoforms described in one entry? UniProt · August 9, 2026. https://www.uniprot.org/help/canonical_and_isoforms
- Shifting the limits in wheat research and breeding using a fully annotated reference genome Science · 2018. https://pubmed.ncbi.nlm.nih.gov/30115783/
- Genome streamlining in a cosmopolitan oceanic bacterium Science · 2005. https://pubmed.ncbi.nlm.nih.gov/16109880/
- Essential genes of a minimal bacterium Proceedings of the National Academy of Sciences · 2006. https://pmc.ncbi.nlm.nih.gov/articles/PMC1324956/
- Small, smaller, smallest: the origins and evolution of ancient dual symbioses in a phloem-feeding insect Genome Biology and Evolution · 2013. https://pmc.ncbi.nlm.nih.gov/articles/PMC3787670/

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.