# What is the smallest protein?

> Chignolin and CLN025 have just 10 amino acids. Compare these designed folds with human RPL41 and MOTS-c, and learn why there is no single smallest protein.

Chignolin and CLN025 show that a designed molecule can display protein-like folding with just 10 amino acids. Among natural molecules, the fruit-fly TAL/pri gene encodes functional 11-amino-acid peptides. In humans, RPL41 is a 25-amino-acid ribosomal protein, while shorter examples include the 16-residue MOTS-c peptide and the 17-residue micropeptide miPEP155.

These numbers answer different questions. A molecule can be short enough to be called a peptide yet have a specific biological function or a stable fold. Identifying the smallest protein therefore requires deciding whether to count designed molecules, natural gene products, processed peptides, or components of larger molecular assemblies.

## Why is there no single smallest protein?

There is no universally accepted minimum number of amino acids in a protein. Length-based naming conventions overlap: UniProt uses its peptide annotation for biologically active products of roughly 40–50 residues or fewer that are processed from larger precursors. This is an annotation rule for a particular class of molecule, not a lower size limit for all proteins.

Here, length means the number of amino acid residues, abbreviated aa, rather than molecular mass or physical diameter. Counts refer to one chain unless a mature multichain molecule is explicitly identified. A short designed fold, a natural developmental regulator, and a component of the ribosome answer different versions of the question.

![Sequence lengths of selected short proteins and peptides, from 10-residue CLN025 to the 147-residue near-infrared tag miRFP670nano.](/images/charts/smallest-proteins-comparison.webp "**Figure 1. Selected examples on one length scale.** The labels distinguish designed folds, natural peptides, and functional tags; these are not equivalent biological units.")

The chart places selected designed folds, natural peptides, enzymes, and fluorescent tags on one length scale. Thymosin beta-4 uses its 44-residue encoded sequence, insulin uses its two mature chains combined, and 4-OT uses one mature subunit. The bars compare residue counts across these stated conventions, not equivalent functional units or contenders for a single record. The individual examples and their sources are discussed below.

The broader term _small protein_ also depends on context. SmProt, a database of short translated products, uses fewer than 100 amino acids as its operational definition. That threshold helps organize a dataset; it does not imply that all shorter sequences share a structure, function, or level of experimental support. For the wider distribution, see [average protein size](/guides/average-protein-size).

## How can a protein fold with only 10 amino acids?

Chignolin and CLN025 form beta-hairpins, in which the short chain turns back on itself and neighboring segments interact. Chignolin was designed using structural patterns from known proteins. Experiments showed that it adopts a defined structure in water and undergoes a cooperative thermal transition, meaning that its folding changes collectively as temperature changes.

CLN025 replaces the two terminal glycines of chignolin with tyrosines. The substitution leaves the residue count unchanged while increasing molecular mass. Its characterization included solution and crystal structures, thermal stability measurements, and an analysis of its folding behavior.

  <caption><strong>Table 1. Chignolin and CLN025 compared.</strong> A kilodalton (kDa) is a unit of molecular mass; the angstrom value describes the resolution of the crystal structure, not the molecule's diameter.</caption>
  <thead>
    <tr>
      <th scope="col">Property</th>
      <th scope="col">Chignolin</th>
      <th scope="col">CLN025</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Sequence</td>
      <td>GYDPETGTWG</td>
      <td>YYDPETGTWY</td>
    </tr>
    <tr>
      <td>Length</td>
      <td>10 amino acids</td>
      <td>10 amino acids</td>
    </tr>
    <tr>
      <td>PDB-reported structure mass</td>
      <td>About 1.08 kDa</td>
      <td>About 1.29 kDa</td>
    </tr>
    <tr>
      <td>Solution NMR structure</td>
      <td>1UAO</td>
      <td>2RVD</td>
    </tr>
    <tr>
      <td>Crystal structure</td>
      <td>Not listed here</td>
      <td>5AWL, resolved at 1.11 Å</td>
    </tr>
  </tbody>

Sequences and structures are reported by Honda and colleagues and in PDB entries 1UAO, 2RVD, and 5AWL.

PDB 1UAO therefore refers to chignolin, not CLN025. Both have the same sequence length, but chignolin is lighter. Their significance is that particular sequences can fold at this scale, not that any chain of 10 amino acids will behave as a protein.

Trp-cage provides a larger comparison. The 20-residue TC5b construct has an experimentally determined NMR structure, and the original study reported constructs that were more than 95% folded in water at physiological pH. These designs help separate the question of how little sequence is needed for folding from the question of how little is needed for biological activity.

## What is the smallest natural protein?

The 11-amino-acid TAL/pri peptides from _Drosophila_ are well-characterized examples of extremely short, functional natural gene products. They are encoded by the tarsal-less/polished rice gene, which contains several short open reading frames. An open reading frame is a stretch of nucleotides that can be translated into a peptide.

Galindo and colleagues showed that these tiny translated products influence gene expression and tissue development, and that one of the short coding units could supply the gene's activity in their experiments. The result established that a functional gene product can be much shorter than traditional protein annotation thresholds.

This evidence concerns biological function. It does not establish an independently stable fold like that of chignolin, or prove that TAL/pri is the shortest functional peptide in every organism. The frequently cited 11-residue example is also a fruit-fly peptide, not a smallest-human-protein record.

## What is the smallest protein in the human body?

There is no undisputed smallest human protein. RPL41 is a 25-amino-acid ribosomal protein, but shorter encoded products include the 16-residue MOTS-c peptide and the 17-residue micropeptide miPEP155. Their biological roles and evidence need to be considered alongside their lengths.

To make this comparison systematic, we analysed all 756 reviewed human entries in UniProtKB release 2026_03 with canonical sequences shorter than 100 amino acids. We screened out incomplete sequences and entries without protein-level evidence, then checked every remaining entry through 31 residues individually. We retained the eight entries below. Our analysis uses existing database annotations and does not establish newly discovered proteins.

![Eight short human proteins and peptides we retained in our UniProt audit: MOTS-c 16, miPEP155 17, Humanin 24, RPL41 25, SHLP2 and PRKCH uORF2 26 each, PTEN MP31 and sarcolipin 31 each.](/images/charts/shortest-human-proteins-audit.webp "**Figure 2. Shortest human entries we retained in our annotation audit.** All have UniProt protein-level evidence; lengths count the listed canonical sequence. Mitochondrial-derived peptides remain subject to production and annotation caveats.")

The chart includes all entries we retained through 31 residues, including ties. Download the [full audit](/data/guides/short-human-proteins/audit.csv), [eight-entry shortlist](/data/guides/short-human-proteins/shortest.csv), or [reproduction files](/data/guides/short-human-proteins/reproduction.zip), which include the frozen UniProt records, methods, review decisions, and Python script.

The 17-residue miPEP155 provides a particularly short example from a nuclear transcript. A 2020 study identified this product of the human MIR155HG transcript and reported evidence of its endogenous expression and a role in antigen presentation, the process by which immune cells display molecular fragments to T cells. It demonstrates that a human nuclear transcript can encode a functional product shorter than RPL41.

RPL41 has the sequence `MRAKWRKKRMRRLKRKRRKMRQRSK` and a reported mass of 3,456 Da. Historically called 60S ribosomal protein L41, it is now named small ribosomal subunit protein eS32 in UniProt, reflecting its structural assignment at the interface of the ribosomal subunits. Its presence in the ribosome establishes a cellular role, but does not demonstrate that the isolated chain forms an independently stable globular protein.

MOTS-c and Humanin need a different qualification. Both have reviewed UniProt entries with protein-level evidence, yet their routes of production remain incompletely understood. UniProt notes that the mitochondrial genetic code would prevent production of the listed MOTS-c sequence within mitochondria. For Humanin, mitochondrial translation would instead yield a shorter, 21-residue peptide; the physiological contribution of related nuclear genes also remains unresolved.

Processing also changes the count for familiar examples. Thymosin beta-4 has 44 encoded residues and 43 after removal of its initial methionine. Human TRH is a chemically modified tripeptide released from a 242-residue precursor. Insulin is translated as a 110-residue precursor; its mature A and B chains contain 21 and 30 residues respectively, or 51 combined. Thus, insulin is small without being the smallest human protein, and a three-residue mature hormone does not imply that the gene's complete translated product is only three residues long. The [insulin amino acid count](/guides/how-many-amino-acids-are-in-insulin) explains this processing in more detail.

### Why not just take the shortest database entry?

The shortest record we retrieved was only two residues long: P0DPR3, a T-cell receptor diversity segment. UniProt marks its sequence ends as non-terminal because it contributes to a larger receptor chain. The four-residue tuftsin entry, P01858, instead describes a peptide released from an immunoglobulin chain. Neither establishes an independently encoded protein of that length.

Other short records require different exclusions. UniProt cautions that the eight-residue urine glycopeptide has not been found in the complete proteome, and that the 11-residue head-activator sequence could not be mapped to the reference genome. Both carry protein-level evidence, showing why that label alone cannot settle a smallest-protein claim.

We used the default sequence in each reviewed human entry, without expanding alternative isoforms or extracting mature peptides from longer precursors. We excluded records marked as fragments or containing non-terminal residues, and required protein-level evidence. Fourteen entries at or below 31 residues survived our screen; after checking their annotations individually, we excluded six processed or unresolved isolated-peptide records. We did not individually certify longer screen survivors. Our shortlist is specific to this database release and these criteria, and cannot rule out shorter products in unreviewed records or the wider literature.

### What evidence exists for even shorter human products?

Sandmann and colleagues reported 221 previously missed human short open reading frames potentially encoding peptides of 3–15 amino acids in 2023. Ribosome profiling, which detects where ribosomes are translating RNA, supported their identification. The authors also reported putative peptide-level evidence for 38 of the 221 candidates, while explicitly noting that false-positive identifications could not be excluded.

This is evidence for a population of candidate short products, not a definitive new smallest human protein. Detecting translation, identifying the resulting molecule, and establishing its function are separate experimental questions.

The distinction remains relevant in larger surveys. In 2026, the TransCODE Consortium reported detectable peptides from about 25% of 7,264 non-canonical open reading frames across 95,520 proteomics experiments. The authors introduced _peptidein_ for translated protein molecules whose functional-protein status remains indeterminate.

The consortium also highlighted a size-related limitation of Human Proteome Project mass-spectrometry guidelines: their usual criteria require two distinct, uniquely mapping peptides of at least nine residues, together covering at least 18 residues. A product shorter than 18 residues cannot meet that coverage requirement. Failing this particular test therefore does not, by itself, show that a very short peptide does not exist.

## How small can enzymes and fluorescent tags be?

4-Oxalocrotonate tautomerase, or 4-OT, from _Pseudomonas putida_ is an established example of a natural enzyme with very short subunits. Its encoded sequence has 63 amino acids; removal of the initiating methionine gives the 62-residue mature chain described in the original study. Six copies assemble into the functional enzyme.

The 62-residue figure describes one subunit, not a complete catalytic assembly. It also describes a bacterial enzyme, so it does not answer the question of the smallest enzyme in the human body. A useful enzyme comparison must specify species, processing, and whether it counts one chain or the full complex.

At a different extreme, a 2019 study reported enzyme-like catalysis by phenylalanine molecules assembled with zinc. The authors described the material as non-proteinaceous. Calling this a one-amino-acid protein enzyme would confuse the building block of a catalytic assembly with a protein chain.

Fluorescent tags illustrate another functional category. NanoFAST contains 98 amino acids and requires an added small-molecule fluorogen to produce fluorescence. Its 2021 paper described it as the shortest known fluorescent or fluorogen-activating protein tag at publication. The 147-residue miRFP670nano, reported in 2019, instead binds biliverdin and was described as the smallest monomeric near-infrared fluorescent protein at that time. These are dated claims within different classes of tag.

## What is the smallest unit of a protein?

An [amino acid](/guides/how-many-amino-acids-are-there) is the basic building block of a protein. Once incorporated into a chain, it is called an amino acid residue. Residues joined through [peptide bonds](/guides/how-many-peptide-bonds-in-a-peptide) form the protein's primary structure, or sequence.

A dipeptide contains two residues joined by one peptide bond. A linear 10-residue chain such as chignolin contains nine backbone peptide bonds. Neither the number of bonds nor the presence of amino acids alone establishes a protein's folding or function.

Glycine is the smallest standard protein-forming amino acid, with a single hydrogen atom as its side chain. Glutathione, by contrast, contains three amino acids but is assembled enzymatically rather than translated as a protein chain. It is a tripeptide metabolite, illustrating why an amino-acid-containing molecule need not be classified as a protein.
