Most proteins are a few hundred amino acids long. A useful rule of thumb is about 300 amino acids for a bacterial protein and 400 amino acids for a eukaryotic protein, equal to roughly 30–50 kilodaltons (kDa).
Physical size is a different measurement. A typical folded, soluble protein is about 3–6 nanometers (nm) across, while elongated proteins and multi-protein complexes can be much larger.
What is the average size of a protein?
A typical protein contains about 300–400 amino acids. In current reviewed reference proteomes, the median is 271 amino acids for E. coli K-12 and 415 amino acids for the human proteome.[1][2]
ProteinIQ calculated the following values from the canonical sequences returned for every reviewed entry in four UniProtKB reference proteomes:
| Reference proteome | Reviewed proteins | Median length | Mean length |
|---|---|---|---|
| E. coli K-12 | 4,403 | 271 aa | 308 aa |
| Budding yeast | 6,067 | 395 aa | 484 aa |
| Arabidopsis thaliana | 16,343 | 391 aa | 453 aa |
| Human | 20,416 | 415 aa | 559 aa |
Source: ProteinIQ calculation from UniProtKB release 2026_02, released June 10, 2026.[1][2][3][4]
The familiar 300-amino-acid bacterial and 400-amino-acid eukaryotic rule remains a good summary.[6] The table is more precise about the dataset: it compares reviewed database entries, not every protein molecule present in a cell.
The human median is higher than the 375-amino-acid figure reported from a 2005 Ensembl protein set.[10] Protein databases change as gene models, start sites, and sequence records are revised, so an average should always name its dataset and date.
How many kilodaltons is an average protein?
The average molecular weight of a protein with 300–400 amino acids is roughly 33–44 kDa. The usual shortcut assigns about 110 daltons to each amino acid residue in a protein chain.[8]
| Sequence length | Approximate mass |
|---|---|
| 100 amino acids | 11 kDa |
| 300 amino acids | 33 kDa |
| 400 amino acids | 44 kDa |
| 500 amino acids | 55 kDa |
| 1,000 amino acids | 110 kDa |
These values are estimates based on 110 Da per residue.[8] Exact mass depends on the amino-acid composition, terminal groups, disulfide bonds, cleavage, and chemical modifications. A molecular-weight calculator uses the actual sequence instead of the shortcut.[9]
For a quick amino-acids-to-kDa conversion, multiply the sequence length by 110 Da and divide by 1,000. To estimate amino acids from kDa, multiply the mass by about 9.1 residues per kDa. A 50 kDa protein is therefore roughly 455 amino acids long, although its exact length depends on its composition.
The residue mass is lower than the average mass of a free amino acid because forming each peptide bond removes the elements of one water molecule. Sequence length counts residues, while the question of how many amino acids exist counts amino-acid types.
How big is a protein in nanometers?
A typical folded, soluble protein is about 3–6 nm in diameter.[6] A compact spherical protein with a mass of 50 kDa has a minimum calculated radius of 2.4 nm, or a diameter of 4.8 nm.[7]
Mass does not determine one exact width. Globular proteins pack into compact shapes, while fibrous proteins, disordered regions, and multi-domain chains can be much longer. Erickson reports a diameter of about 5 nm for hemoglobin, while rod-shaped fibrinogen is about 46 nm long.[7]
An unfolded chain is longer again because contour length follows the backbone rather than the packed volume. Amino-acid count, molecular mass, and physical diameter should not be used interchangeably.
How large are GAPDH, beta-actin, and other common proteins?
Common protein molecular weights range from 26.9 kDa for green fluorescent protein (GFP) to 158.4 kDa for Streptococcus pyogenes Cas9 in the examples below. Human GAPDH is 36.1 kDa, beta-actin is 41.7 kDa, and alpha-tubulin is 50.2 kDa.[5]
| Protein record | UniProt accession | Length | Sequence mass |
|---|---|---|---|
| Green fluorescent protein | P42212 | 238 aa | 26.9 kDa |
| Human GAPDH | P04406 | 335 aa | 36.1 kDa |
| Human beta-actin | P60709 | 375 aa | 41.7 kDa |
| SARS-CoV-2 nucleocapsid | P0DTC9 | 419 aa | 45.6 kDa |
| Human alpha-tubulin 1B | P68363 | 451 aa | 50.2 kDa |
| Human albumin precursor | P02768 | 609 aa | 69.4 kDa |
| SARS-CoV-2 spike | P0DTC2 | 1,273 aa | 141.2 kDa |
| S. pyogenes Cas9 | Q99ZW2 | 1,368 aa | 158.4 kDa |
The values are the unmodified sequence lengths and masses in UniProtKB release 2026_02.[5] Albumin is listed as its translated precursor; the mature circulating chain is shorter after signal peptide and propeptide removal.
These sequence masses are good reference points for a Western blot, but an observed band can shift because of cleavage, glycosylation, phosphorylation, other modifications, or unusual migration in the gel. A molecular-weight marker is a calibration standard, not a measure of the average protein size.
Why do reported average protein sizes differ?
Reported averages differ because researchers may count different sequences and summarize them differently. The median describes the midpoint protein; the arithmetic mean is pulled upward by rare giants such as titin.
That effect is visible in the current human data. The median reviewed human sequence is 415 amino acids, but the mean is 559 because the distribution has a long upper tail.[1] At the small end, short peptides and microproteins can also change the result when a database includes or excludes them.
The unit being counted matters too. A database can report a translated precursor, a mature cleaved chain, an isoform, or one subunit of a larger complex. An abundance-weighted average asks yet another question by giving common cellular proteins more weight than rare ones.[6]
For the UniProtKB table above, ProteinIQ downloaded the reviewed canonical entries in each named reference proteome on August 9, 2026, then calculated the median and arithmetic mean of the sequence-length field. The values will change as UniProt revises the underlying records.
Sources▼
- Reviewed human reference proteome lengths (UniProtKB release 2026_02) UniProt Consortium · August 9, 2026. https://rest.uniprot.org/uniprotkb/stream?query=proteome%3AUP000005640%20AND%20reviewed%3Atrue&format=tsv&fields=accession%2Clength
- Reviewed E. coli K-12 reference proteome lengths (UniProtKB release 2026_02) UniProt Consortium · August 9, 2026. https://rest.uniprot.org/uniprotkb/stream?query=proteome%3AUP000000625%20AND%20reviewed%3Atrue&format=tsv&fields=accession%2Clength
- Reviewed budding yeast reference proteome lengths (UniProtKB release 2026_02) UniProt Consortium · August 9, 2026. https://rest.uniprot.org/uniprotkb/stream?query=proteome%3AUP000002311%20AND%20reviewed%3Atrue&format=tsv&fields=accession%2Clength
- Reviewed Arabidopsis reference proteome lengths (UniProtKB release 2026_02) UniProt Consortium · August 9, 2026. https://rest.uniprot.org/uniprotkb/stream?query=proteome%3AUP000006548%20AND%20reviewed%3Atrue&format=tsv&fields=accession%2Clength
- UniProtKB records for GAPDH, beta-actin, GFP, alpha-tubulin, albumin, Cas9, and SARS-CoV-2 proteins UniProt Consortium · August 10, 2026. https://rest.uniprot.org/uniprotkb/search?query=%28accession%3AP04406%20OR%20accession%3AP60709%20OR%20accession%3AP42212%20OR%20accession%3AP68363%20OR%20accession%3AP02768%20OR%20accession%3AQ99ZW2%20OR%20accession%3AP0DTC2%20OR%20accession%3AP0DTC9%29&format=tsv&fields=accession%2Cid%2Cprotein_name%2Clength%2Cmass%2Corganism_name
- How big is the average protein? Cell Biology by the Numbers · 2015. https://book.bionumbers.org/how-big-is-the-average-protein/
- Size and shape of protein molecules at the nanometer level determined by sedimentation, gel filtration, and electron microscopy Biological Procedures Online · 2009. https://pmc.ncbi.nlm.nih.gov/articles/PMC3055910/
- Biological macromolecules and amino acids eCampusOntario Pressbooks · 2021. https://ecampusontario.pressbooks.pub/bioc2580/chapter/bioc2580-lecture-1-biological-macromolecules-amino-acids/
- ProtParam documentation SIB Swiss Institute of Bioinformatics · August 9, 2026. https://web.expasy.org/protparam/protparam-doc.html
- Protein length in eukaryotic and prokaryotic proteomes Nucleic Acids Research · 2005. https://pmc.ncbi.nlm.nih.gov/articles/PMC1150220/

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.