ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Proteins

What is the smallest protein?

September 19, 2026·Matic Broz, PhD
Conceptual illustration comparing a long protein ribbon with a magnified compact protein fold.

TL;DR

Chignolin and CLN025 are among the smallest designed protein-like molecules, each containing just 10 amino acids and forming a defined fold. There is no universally accepted smallest protein: natural and human examples depend on the definition. Human RPL41 contains 25 amino acids, while the bioactive peptide MOTS-c contains only 16.

Chignolin and CLN025 show that a designed molecule can display protein-like folding with just 10 amino acids. Among natural molecules, the fruit-fly TAL/pri gene encodes functional 11-amino-acid peptides. In humans, RPL41 is a 25-amino-acid ribosomal protein, while shorter examples include the 16-residue MOTS-c peptide and the 17-residue micropeptide miPEP155.

These numbers answer different questions. A molecule can be short enough to be called a peptide yet have a specific biological function or a stable fold. Identifying the smallest protein therefore requires deciding whether to count designed molecules, natural gene products, processed peptides, or components of larger molecular assemblies.

Why is there no single smallest protein?

There is no universally accepted minimum number of amino acids in a protein. Length-based naming conventions overlap: UniProt uses its peptide annotation for biologically active products of roughly 40–50 residues or fewer that are processed from larger precursors. This is an annotation rule for a particular class of molecule, not a lower size limit for all proteins.[1]

Here, length means the number of amino acid residues, abbreviated aa, rather than molecular mass or physical diameter. Counts refer to one chain unless a mature multichain molecule is explicitly identified. A short designed fold, a natural developmental regulator, and a component of the ribosome answer different versions of the question.

Figure 1. Selected examples on one length scale. The labels distinguish designed folds, natural peptides, and functional tags; these are not equivalent biological units. Reuse under CC BY 4.0.

The chart places selected designed folds, natural peptides, enzymes, and fluorescent tags on one length scale. Thymosin beta-4 uses its 44-residue encoded sequence, insulin uses its two mature chains combined, and 4-OT uses one mature subunit. The bars compare residue counts across these stated conventions, not equivalent functional units or contenders for a single record.[2][3][4] The individual examples and their sources are discussed below.

The broader term small protein also depends on context. SmProt, a database of short translated products, uses fewer than 100 amino acids as its operational definition. That threshold helps organize a dataset; it does not imply that all shorter sequences share a structure, function, or level of experimental support.[5] For the wider distribution, see average protein size.

How can a protein fold with only 10 amino acids?

Chignolin and CLN025 form beta-hairpins, in which the short chain turns back on itself and neighboring segments interact. Chignolin was designed using structural patterns from known proteins. Experiments showed that it adopts a defined structure in water and undergoes a cooperative thermal transition, meaning that its folding changes collectively as temperature changes.[6]

CLN025 replaces the two terminal glycines of chignolin with tyrosines. The substitution leaves the residue count unchanged while increasing molecular mass. Its characterization included solution and crystal structures, thermal stability measurements, and an analysis of its folding behavior.[7]

PropertyChignolinCLN025
SequenceGYDPETGTWGYYDPETGTWY
Length10 amino acids10 amino acids
PDB-reported structure massAbout 1.08 kDaAbout 1.29 kDa
Solution NMR structure1UAO2RVD
Crystal structureNot listed here5AWL, resolved at 1.11 Å
Table 1. Chignolin and CLN025 compared. A kilodalton (kDa) is a unit of molecular mass; the angstrom value describes the resolution of the crystal structure, not the molecule's diameter.

Sequences and structures are reported by Honda and colleagues and in PDB entries 1UAO, 2RVD, and 5AWL.[6][7][8][9]

PDB 1UAO therefore refers to chignolin, not CLN025. Both have the same sequence length, but chignolin is lighter. Their significance is that particular sequences can fold at this scale, not that any chain of 10 amino acids will behave as a protein.

Trp-cage provides a larger comparison. The 20-residue TC5b construct has an experimentally determined NMR structure, and the original study reported constructs that were more than 95% folded in water at physiological pH.[10] These designs help separate the question of how little sequence is needed for folding from the question of how little is needed for biological activity.

What is the smallest natural protein?

The 11-amino-acid TAL/pri peptides from Drosophila are well-characterized examples of extremely short, functional natural gene products. They are encoded by the tarsal-less/polished rice gene, which contains several short open reading frames. An open reading frame is a stretch of nucleotides that can be translated into a peptide.[11]

Galindo and colleagues showed that these tiny translated products influence gene expression and tissue development, and that one of the short coding units could supply the gene's activity in their experiments. The result established that a functional gene product can be much shorter than traditional protein annotation thresholds.[11]

This evidence concerns biological function. It does not establish an independently stable fold like that of chignolin, or prove that TAL/pri is the shortest functional peptide in every organism. The frequently cited 11-residue example is also a fruit-fly peptide, not a smallest-human-protein record.

What is the smallest protein in the human body?

There is no undisputed smallest human protein. RPL41 is a 25-amino-acid ribosomal protein, but shorter encoded products include the 16-residue MOTS-c peptide and the 17-residue micropeptide miPEP155. Their biological roles and evidence need to be considered alongside their lengths.[12][13][14]

To make this comparison systematic, we analysed all 756 reviewed human entries in UniProtKB release 2026_03 with canonical sequences shorter than 100 amino acids. We screened out incomplete sequences and entries without protein-level evidence, then checked every remaining entry through 31 residues individually. We retained the eight entries below. Our analysis uses existing database annotations and does not establish newly discovered proteins.[15]

Figure 2. Shortest human entries we retained in our annotation audit. All have UniProt protein-level evidence; lengths count the listed canonical sequence. Mitochondrial-derived peptides remain subject to production and annotation caveats. Reuse under CC BY 4.0.

The chart includes all entries we retained through 31 residues, including ties. Download the full audit, eight-entry shortlist, or reproduction files, which include the frozen UniProt records, methods, review decisions, and Python script.[15]

The 17-residue miPEP155 provides a particularly short example from a nuclear transcript. A 2020 study identified this product of the human MIR155HG transcript and reported evidence of its endogenous expression and a role in antigen presentation, the process by which immune cells display molecular fragments to T cells. It demonstrates that a human nuclear transcript can encode a functional product shorter than RPL41.[14]

RPL41 has the sequence MRAKWRKKRMRRLKRKRRKMRQRSK and a reported mass of 3,456 Da. Historically called 60S ribosomal protein L41, it is now named small ribosomal subunit protein eS32 in UniProt, reflecting its structural assignment at the interface of the ribosomal subunits. Its presence in the ribosome establishes a cellular role, but does not demonstrate that the isolated chain forms an independently stable globular protein.[12]

MOTS-c and Humanin need a different qualification. Both have reviewed UniProt entries with protein-level evidence, yet their routes of production remain incompletely understood. UniProt notes that the mitochondrial genetic code would prevent production of the listed MOTS-c sequence within mitochondria. For Humanin, mitochondrial translation would instead yield a shorter, 21-residue peptide; the physiological contribution of related nuclear genes also remains unresolved.[13][16]

Processing also changes the count for familiar examples. Thymosin beta-4 has 44 encoded residues and 43 after removal of its initial methionine.[2] Human TRH is a chemically modified tripeptide released from a 242-residue precursor. Insulin is translated as a 110-residue precursor; its mature A and B chains contain 21 and 30 residues respectively, or 51 combined. Thus, insulin is small without being the smallest human protein, and a three-residue mature hormone does not imply that the gene's complete translated product is only three residues long.[17][3] The insulin amino acid count explains this processing in more detail.

Why not just take the shortest database entry?

The shortest record we retrieved was only two residues long: P0DPR3, a T-cell receptor diversity segment. UniProt marks its sequence ends as non-terminal because it contributes to a larger receptor chain. The four-residue tuftsin entry, P01858, instead describes a peptide released from an immunoglobulin chain. Neither establishes an independently encoded protein of that length.[15]

Other short records require different exclusions. UniProt cautions that the eight-residue urine glycopeptide has not been found in the complete proteome, and that the 11-residue head-activator sequence could not be mapped to the reference genome. Both carry protein-level evidence, showing why that label alone cannot settle a smallest-protein claim.[15]

We used the default sequence in each reviewed human entry, without expanding alternative isoforms or extracting mature peptides from longer precursors. We excluded records marked as fragments or containing non-terminal residues, and required protein-level evidence. Fourteen entries at or below 31 residues survived our screen; after checking their annotations individually, we excluded six processed or unresolved isolated-peptide records. We did not individually certify longer screen survivors. Our shortlist is specific to this database release and these criteria, and cannot rule out shorter products in unreviewed records or the wider literature.[15]

What evidence exists for even shorter human products?

Sandmann and colleagues reported 221 previously missed human short open reading frames potentially encoding peptides of 3–15 amino acids in 2023. Ribosome profiling, which detects where ribosomes are translating RNA, supported their identification. The authors also reported putative peptide-level evidence for 38 of the 221 candidates, while explicitly noting that false-positive identifications could not be excluded.[18]

This is evidence for a population of candidate short products, not a definitive new smallest human protein. Detecting translation, identifying the resulting molecule, and establishing its function are separate experimental questions.

The distinction remains relevant in larger surveys. In 2026, the TransCODE Consortium reported detectable peptides from about 25% of 7,264 non-canonical open reading frames across 95,520 proteomics experiments. The authors introduced peptidein for translated protein molecules whose functional-protein status remains indeterminate.[19]

The consortium also highlighted a size-related limitation of Human Proteome Project mass-spectrometry guidelines: their usual criteria require two distinct, uniquely mapping peptides of at least nine residues, together covering at least 18 residues. A product shorter than 18 residues cannot meet that coverage requirement. Failing this particular test therefore does not, by itself, show that a very short peptide does not exist.[19]

How small can enzymes and fluorescent tags be?

4-Oxalocrotonate tautomerase, or 4-OT, from Pseudomonas putida is an established example of a natural enzyme with very short subunits. Its encoded sequence has 63 amino acids; removal of the initiating methionine gives the 62-residue mature chain described in the original study. Six copies assemble into the functional enzyme.[20][4]

The 62-residue figure describes one subunit, not a complete catalytic assembly. It also describes a bacterial enzyme, so it does not answer the question of the smallest enzyme in the human body. A useful enzyme comparison must specify species, processing, and whether it counts one chain or the full complex.

At a different extreme, a 2019 study reported enzyme-like catalysis by phenylalanine molecules assembled with zinc. The authors described the material as non-proteinaceous. Calling this a one-amino-acid protein enzyme would confuse the building block of a catalytic assembly with a protein chain.[21]

Fluorescent tags illustrate another functional category. NanoFAST contains 98 amino acids and requires an added small-molecule fluorogen to produce fluorescence. Its 2021 paper described it as the shortest known fluorescent or fluorogen-activating protein tag at publication. The 147-residue miRFP670nano, reported in 2019, instead binds biliverdin and was described as the smallest monomeric near-infrared fluorescent protein at that time. These are dated claims within different classes of tag.[22][23]

What is the smallest unit of a protein?

An amino acid is the basic building block of a protein. Once incorporated into a chain, it is called an amino acid residue. Residues joined through peptide bonds form the protein's primary structure, or sequence.[24]

A dipeptide contains two residues joined by one peptide bond. A linear 10-residue chain such as chignolin contains nine backbone peptide bonds. Neither the number of bonds nor the presence of amino acids alone establishes a protein's folding or function.

Glycine is the smallest standard protein-forming amino acid, with a single hydrogen atom as its side chain.[25] Glutathione, by contrast, contains three amino acids but is assembled enzymatically rather than translated as a protein chain. It is a tripeptide metabolite, illustrating why an amino-acid-containing molecule need not be classified as a protein.[26]

Sources26
  1. Peptide annotation definition

    UniProt · September 19, 2026

  2. Thymosin beta-4, UniProtKB P62328

    UniProt · September 19, 2026

  3. Insulin, UniProtKB P01308

    UniProt · September 19, 2026

  4. 2-hydroxymuconate tautomerase, UniProtKB Q01468

    UniProt · September 19, 2026

  5. SmProt: A Reliable Repository with Comprehensive Annotation of Small Proteins Identified from Ribosome Profiling

    Genomics, Proteomics & Bioinformatics · 2021

  6. 1UAO: NMR structure of designed protein, chignolin, consisting of only ten amino acids

    RCSB Protein Data Bank · September 19, 2026

  7. Crystal Structure of a Ten-Amino Acid Protein

    Journal of the American Chemical Society · 2008

  8. 2RVD: NMR structure of a mutant of chignolin, CLN025

    RCSB Protein Data Bank · September 19, 2026

  9. 5AWL: Crystal structure of a mutant of chignolin, CLN025

    RCSB Protein Data Bank · September 19, 2026

  10. 1L2Y: NMR structure of Trp-cage miniprotein construct TC5b

    RCSB Protein Data Bank · September 19, 2026

  11. Peptides encoded by short ORFs control development and define a new eukaryotic gene family

    PLOS Biology · 2007

  12. Small ribosomal subunit protein eS32, UniProtKB P62945

    UniProt · September 19, 2026

  13. Mitochondrial-derived peptide MOTS-c, UniProtKB A0A0C5B5G6

    UniProt · September 19, 2026

  14. A micropeptide encoded by lncRNA MIR155HG suppresses autoimmune inflammation via modulating antigen presentation

    Science Advances · 2020

  15. Short human proteins: UniProt 2026_03 annotation audit, dataset and methods

    ProteinIQ; source data from UniProt · September 19, 2026

  16. Humanin, UniProtKB Q8IVG9

    UniProt · September 19, 2026

  17. Pro-thyrotropin-releasing hormone, UniProtKB P20396

    UniProt · September 19, 2026

  18. Evolutionary origins and interactomes of human, young microproteins and small peptides translated from short open reading frames

    Molecular Cell · 2023

  19. Expanding the human proteome with microproteins and peptideins

    Nature · 2026

  20. 4-Oxalocrotonate tautomerase, an enzyme composed of 62 amino acid residues per monomer

    Journal of Biological Chemistry · 1992

  21. Non-proteinaceous hydrolase comprised of a phenylalanine metallo-supramolecular amyloid-like structure

    Nature Catalysis · 2019

  22. NanoFAST: structure-based design of a small fluorogen-activating protein with only 98 amino acids

    Chemical Science · 2021

  23. Smallest near-infrared fluorescent protein evolved from cyanobacteriochrome as versatile tag for spectral multiplexing

    Nature Communications · 2019

  24. Biochemistry, Primary Protein Structure

    StatPearls, NCBI Bookshelf · September 19, 2026

  25. Glycine, ChEBI 15428

    ChEBI · September 19, 2026

  26. Glutathione: Overview of its protective roles, measurement, and biosynthesis

    Molecular Aspects of Medicine · 2009

Cite this article

Broz, M. (2026, September 19). What is the smallest protein? ProteinIQ. https://proteiniq.io/guides/smallest-protein

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “What is the smallest protein?” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
December 29, 2025
Updated
September 19, 2026

Related guides

Browse all guides
Open peptide chain with one connecting peptide bond highlighted.

Proteins · July 27, 2026

How many peptide bonds are in a peptide?

A single linear peptide chain has one fewer peptide bond than amino acid residues. See the counts for dipeptides through hexapeptides, 100-residue chains, insulin, and cyclic peptides.

Illustrated peptide chain and sample vial with labels for length, purity, scale and modifications.

Proteins · October 5, 2026

How much does peptide synthesis cost?

Custom peptides list at about $2 to $4.50 per amino acid crude and about $12 per amino acid at more than 95% purity. See prices by purity, length, scale and modification.

Human silhouette beside illustrations of collagen, a folded enzyme, and a Y-shaped antibody.

Proteins · September 29, 2026

How many types of proteins are in the human body?

The human body makes nearly 20,000 distinct proteins, grouped into about seven functional types and dozens of database classes. Collagen alone has 28 types, and all are built from 20 standard amino acids.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing