Leucine is the most abundant standard amino acid in reviewed protein sequences. ProteinIQ counted the current UniProtKB/Swiss-Prot release and found that leucine makes up 20,152,734 of 208,897,834 standard amino-acid residues, or 9.65%.
That percentage describes residue frequency in an all-species protein-sequence database. It does not describe codon frequency, dietary intake, free amino-acid concentration, or every organism's proteome.
Which amino acid is most abundant in proteins?
Leucine is the most abundant standard amino acid in current UniProtKB/Swiss-Prot, at 9.65% of standard residues in release 2026_02.[1][2]
Swiss-Prot is useful for this question because it is the reviewed, expert-curated section of UniProtKB, not a raw dump of every predicted protein sequence. The 2026_02 release contains 575,503 sequence entries and 208,906,902 amino-acid symbols; ProteinIQ's calculation counted the standard 20 amino acids separately from ambiguous and rare symbols.[1][4]
The top six residues in the current release are leucine, alanine, glycine, valine, glutamic acid, and serine. Tryptophan is the least common standard amino acid at 1.11%. The full distribution shows that leucine is almost nine times as frequent as tryptophan.
| Rank | Amino acid | 3-letter | One-letter | Count | Frequency |
|---|---|---|---|---|---|
| 1 | Leucine | Leu | L | 20,152,734 | 9.65% |
| 2 | Alanine | Ala | A | 17,247,358 | 8.26% |
| 3 | Glycine | Gly | G | 14,774,258 | 7.07% |
| 4 | Valine | Val | V | 14,318,621 | 6.85% |
| 5 | Glutamic acid | Glu | E | 14,025,555 | 6.71% |
| 6 | Serine | Ser | S | 13,925,029 | 6.67% |
| 7 | Isoleucine | Ile | I | 12,333,032 | 5.90% |
| 8 | Lysine | Lys | K | 12,106,912 | 5.80% |
| 9 | Arginine | Arg | R | 11,548,011 | 5.53% |
| 10 | Aspartic acid | Asp | D | 11,411,936 | 5.46% |
| 11 | Threonine | Thr | T | 11,213,040 | 5.37% |
| 12 | Proline | Pro | P | 9,926,740 | 4.75% |
| 13 | Asparagine | Asn | N | 8,491,928 | 4.07% |
| 14 | Glutamine | Gln | Q | 8,217,450 | 3.93% |
| 15 | Phenylalanine | Phe | F | 8,082,394 | 3.87% |
| 16 | Tyrosine | Tyr | Y | 6,110,160 | 2.92% |
| 17 | Methionine | Met | M | 5,037,584 | 2.41% |
| 18 | Histidine | His | H | 4,761,702 | 2.28% |
| 19 | Cysteine | Cys | C | 2,902,631 | 1.39% |
| 20 | Tryptophan | Trp | W | 2,310,759 | 1.11% |
Source: ProteinIQ calculation from the UniProtKB/Swiss-Prot release 2026_02 FASTA file, checked against UniProt's release statistics and MD5 metadata.[1][2][3]
An independent analysis also found leucine to be the most abundant amino acid in both Swiss-Prot and TrEMBL, with tryptophan and cysteine among the least abundant.[10]
This database-wide protein statistic is not a biological constant. A 2024 Scientific Reports study of 5,590 proteomes found that amino-acid usage differs across domains of life, even though only a few amino acids tend to dominate the most-used and least-used ranks.[9] Individual proteins, organisms, tissues, and protein families can therefore differ sharply from the Swiss-Prot average.
Why is glycine the most common amino acid in collagen?
Glycine is the most common amino acid in collagen because collagen repeats a Gly-X-Y pattern, so every third residue is glycine.
That gives collagen a simple derived glycine share: one residue in three, or about 33.3%. NCBI Bookshelf's collagen synthesis chapter describes the primary collagen sequence as glycine-proline-X or glycine-X-hydroxyproline and states that every third amino acid is glycine.[5]
This is also why "the most common amino acid in the human body" is a more ambiguous question than it looks. If the question means a broad protein-sequence database, the answer is leucine. If it means collagen-rich connective tissue, the answer is glycine. If it means all body material by mass, amino acids are not the right denominator because the body also contains water, lipids, minerals, nucleic acids, carbohydrates, and many small molecules.
Does amino acid abundance depend on what you count?
Yes. Amino acid abundance can mean residue frequency in a sequence database, composition of one organism's proteome, codon usage, free amino-acid concentration, or dietary supply. Those measurements are not interchangeable.
| Measurement | What it counts | Does the 9.65% figure apply? |
|---|---|---|
| Swiss-Prot sequence abundance | Standard residues across reviewed proteins from all species | Yes, for release 2026_02 |
| Human proteome composition | Residues in a human-only protein set | No; it needs a human-only calculation |
| Codon frequency | DNA or RNA triplets in coding sequences | No; several codons can encode the same amino acid |
| Free amino-acid concentration | Unbound amino acids in blood, cells, or another sample | No; these are not residues inside proteins |
| Dietary or limiting amino acids | Amino-acid supply relative to an organism's nutritional needs | No; the answer depends on the diet and organism |
The sequence and genetic-code distinctions follow the Swiss-Prot dataset and cross-proteome analysis; the dietary distinction follows nutrition and swine-feed references.[1][6][7][8][10]
For one sequence, the amino acid composition tool counts each residue directly. A human proteome calculation would instead filter the input to a defined human protein set.
How did ProteinIQ calculate this 2026 table?
ProteinIQ downloaded the UniProtKB/Swiss-Prot release 2026_02 FASTA file, verified the file against UniProt's release metadata, and counted residues from the sequence lines.
The calculation used scripts/swissprot_amino_acid_frequencies.py. The generated data file is data/generated/swissprot-amino-acid-frequencies-2026_02.json, and the generated Markdown table is data/generated/swissprot-amino-acid-frequencies-2026_02.md.
The denominator for the frequency table is the 208,897,834 standard residues among the 20 common amino acids. The FASTA file also contained 9,068 non-standard or ambiguous symbols: B 276, O 29, U 331, X 8,183, and Z 249. Those symbols are reported separately and excluded from the 20-amino-acid frequency percentages.
The script counted 575,503 FASTA entries and 208,906,902 total sequence symbols, matching the official UniProtKB/Swiss-Prot release statistics. The FASTA MD5 recorded in UniProt's release metadata was 797dad11a33b1b58e3c140649a74d6b6.[1][3]
Sources▼
- UniProtKB/Swiss-Prot Release 2026_02 statistics UniProtKB/Swiss-Prot · August 9, 2026. https://web.expasy.org/docs/relnotes/relstat.html
- UniProtKB/Swiss-Prot release 2026_02 FASTA UniProt · August 9, 2026. https://ftp.uniprot.org/pub/databases/uniprot/current_release/knowledgebase/complete/uniprot_sprot.fasta.gz
- UniProt release 2026_02 metadata and checksums UniProt · August 9, 2026. https://ftp.uniprot.org/pub/databases/uniprot/current_release/knowledgebase/complete/RELEASE.metalink
- UniProtKB SIB Swiss Institute of Bioinformatics · August 9, 2026. https://www.expasy.org/resources/uniprotkb
- Biochemistry, Collagen Synthesis StatPearls, NCBI Bookshelf · Updated September 4, 2023. https://www.ncbi.nlm.nih.gov/books/NBK507709/
- Protein and Amino Acids Recommended Dietary Allowances, NCBI Bookshelf · 1989. https://www.ncbi.nlm.nih.gov/books/NBK234922/
- Formulating farm-specific swine diets University of Minnesota Extension · Reviewed in 2024. https://extension.umn.edu/swine-nutrition/formulating-farm-specific-swine-diets
- Protein and Amino Acid Sources for Swine Diets Pork Information Gateway · 2010. https://porkgateway.org/resource/protein-and-amino-acid-sources-for-swine-diets/
- Differential amino acid usage leads to ubiquitous edge effect in proteomes across domains of life Scientific Reports · 2024. https://www.nature.com/articles/s41598-024-77319-4
- Amino Acid Metabolism Conflicts with Protein Diversity Molecular Biology and Evolution · 2014. https://pmc.ncbi.nlm.nih.gov/articles/PMC4209132/

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.