TL;DR
- GENCODE v50 lists 19,442 human protein-coding genes and 172,117 distinct translations.
- Sequence variation, protein processing, and chemical modification expand those translations into hundreds of thousands or millions of proteoforms.
- A 2005 Berkeley Lab estimate put the number of protein types across life at about 50 billion, but this was an extrapolation rather than a catalogue.
- A 2023 HeLa map quantified 12,653 canonical proteins whose copy numbers added up to 3.36 billion molecules.
- UniProtKB contained 149,810,139 protein records across species in release 2026_02.
How many proteins are there?
Humans have about 20,000 reference protein types, based on 19,442 protein-coding genes in GENCODE v50. Alternative splicing expands those genes into 172,117 distinct translated sequences in current annotation. Sequence variation, cleavage, and chemical modification expand that set into hundreds of thousands or millions of distinct protein forms.[1][3]
The roughly 70,000 human proteins still quoted in some articles came from an older Ensembl count cited in the 2018 proteoform review. It should not be treated as today's upper bound for translated sequences.[3]
Across all life on Earth, no one has counted every protein type. A 2005 Berkeley Lab estimate put the total at about 50 billion, based on the estimated number of living species. It remains an extrapolation, while UniProtKB provides a count of sequences that have been catalogued: 149,810,139 records in release 2026_02.[5][2]
| What is being counted? | Current answer | What the figure means |
|---|---|---|
| Human protein-coding genes | 19,442 | Genes that encode proteins in GENCODE v50 |
| Distinct human translations | 172,117 | Annotated amino acid sequences translated from human transcripts |
| Human proteoforms | Hundreds of thousands to millions | Protein forms created by splicing, sequence variation, processing, and modification |
| Protein types across life | About 50 billion | A 2005 extrapolation based on the estimated number of species |
| Catalogued protein records | 149,810,139 | UniProtKB records across species in release 2026_02 |
There is no accepted final count for human proteoforms. A proteoform is a specific molecular form of a protein, and the set present changes among tissues, cells, and biological conditions. Published estimates therefore depend on whether they count one representative protein per gene, translated isoforms, or modified proteoforms.[3] The human proteome guide examines those definitions in more detail.
How many different proteins are in a human cell?
The HeLa map contained 12,653 quantified canonical proteins. Together, their copy numbers added up to 3.36 billion molecules.[4] This is the clearest way to separate the two common meanings of protein count: 12,653 refers to protein identities in the map, while 3.36 billion refers to all their physical copies.
No fixed number applies to every human cell. A liver cell, neuron, and immune cell express different sets of proteins, and expression changes with age, health, and the cell's current activity. Proteomics also has difficulty detecting proteins present at very low abundance.
The HeLa study focused on canonical proteins and covered about 60% of predicted human protein-coding genes. Its 12,653 figure is a measured reference for one cell line and the protein forms detectable by the methods used.[4]
How many protein molecules are in a cell?
The total depends heavily on cell size. A 2023 analysis estimated about 5.85 million protein molecules in an E. coli cell, 81.63 million in a budding yeast cell, and 3.36 billion in a HeLa cell.[4]
| Cell | Estimated protein molecules | Quantified canonical proteins | Molecules per µm³ |
|---|---|---|---|
| E. coli | 5,852,319 | 3,852 | 2.69 million |
| Budding yeast | 81,627,580 | 4,680 | 2.03 million |
| HeLa | 3,360,824,528 | 12,653 | 1.53 million |
These figures are estimates from integrated proteome maps. The researchers combined quantitative proteomics datasets and normalized them to estimates of total protein mass per cell. Before normalization, published totals for the same cell type differed by as much as tenfold.[4]
The estimates apply to specific cells and growth conditions. A larger cell usually contains more protein molecules, while the amount of each individual protein can range from a few copies to millions.
How many protein sequences have been catalogued?
UniProtKB contained 149,810,139 protein records in release 2026_02. Of these, 575,503 were manually reviewed Swiss-Prot records and 149,234,636 were unreviewed TrEMBL records.[2]
| Database | Records | What they represent |
|---|---|---|
| UniProtKB | 149,810,139 | All reviewed and unreviewed protein records |
| Swiss-Prot | 575,503 | Manually reviewed records |
| TrEMBL | 149,234,636 | Computationally annotated, unreviewed records |
A database record describes a sequence annotation. Databases contain sequences from many species, strains, and variants, so one record does not correspond to one physical molecule and records are not always unique protein types. The totals change with every release.
How can a small set of amino acids make so many proteins?
Most protein sequences are built from 20 standard amino acids joined in different orders and chain lengths. Genes can also encode selenocysteine and pyrrolysine, bringing the known genetically encoded set to 22; pyrrolysine has been found in some bacteria and archaea, but not in humans.[6][7]
Protein structure is described at four levels. Primary structure is the amino acid sequence. Secondary structure describes local helices and sheets, tertiary structure is the folded three-dimensional chain, and quaternary structure describes interactions among multiple protein chains.[6]
Why do protein counts differ so much?
Protein counts move through several levels. Genes describe what an organism can encode. Translations describe the amino acid sequences produced from annotated transcripts. Proteoforms add sequence variation, processing, and chemical modification. A cell then contains many physical copies of the proteins it expresses.
Database records introduce another denominator by collecting sequences across organisms. Adding these figures together would mix biological copies, sequence identities, genomic annotations, and database entries.
Methodology
The per-cell figures come from Dolgalev and colleagues, who integrated published quantitative proteomics datasets for E. coli, budding yeast, and HeLa cells. They normalized protein copy numbers using literature estimates of total protein mass and typical cell volume, then summed the averaged copy numbers in each integrated proteome.[4] The article reports their exact integrated totals in the table and rounds them in the prose.
The human gene, transcript, and translation counts come directly from GENCODE v50.[1] The database counts are the UniProtKB release 2026_02 totals accessed on August 9, 2026.[2] These sources use different counting rules, so no calculation combines them into a single total.
Sources▼
- Human release statistics (v50) GENCODE · August 9, 2026. https://www.gencodegenes.org/human/stats.html
- UniProtKB search results, release 2026_02 UniProt · August 9, 2026. https://www.uniprot.org/uniprotkb?query=*
- How many human proteoforms are there? Nature Chemical Biology · 2018. https://pmc.ncbi.nlm.nih.gov/articles/PMC5837046/
- Estimating Total Quantitative Protein Content in Escherichia coli, Saccharomyces cerevisiae, and HeLa Cells International Journal of Molecular Sciences · 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC9916689/
- A New Guide to Exploring the Protein Universe Lawrence Berkeley National Laboratory · 2005. https://www2.lbl.gov/Science-Articles/Archive/sabl/2005/March/02-protein-universe.html
- Proteins by the Numbers National Institute of General Medical Sciences · 2025. https://nigms.nih.gov/biobeat/2025/01/proteins-by-the-numbers
- What are proteins and what do they do? MedlinePlus Genetics · August 23, 2026. https://medlineplus.gov/genetics/understanding/howgeneswork/protein/

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.