# Short human proteins: UniProt annotation audit

We analysed UniProtKB release 2026_03 (September 2, 2026), which we retrieved
on September 19, 2026. Our analysis uses existing annotations; we do not claim
experimental discovery or identification of the smallest human protein.

## Scope and method

We queried all reviewed human entries whose default canonical sequence has
1–99 amino acid residues, using the query in `sources.json`. We fetched both
result pages and checked the release and reported total for consistency. We
counted records by accession without expanding alternative isoforms or grouping
by unique sequence or gene. We retrieved 756 records.

We screened out entries with a Fragment flag, a Non-terminal
residue feature, or a protein-existence level other than 1 (evidence at protein
level). Multiple reasons can apply to one entry. We excluded 230 records at this stage.
Absence of a Fragment flag is not sufficient to establish a complete translation:
immune-receptor segments can carry Non-terminal residue annotations instead.

We then read every surviving entry of 31 residues or fewer individually. We
recorded our 14 decisions in `review.json`, including entry versions and reasons.
We excluded six entries: tuftsin is processed from a longer immunoglobulin
chain; five other isolated-peptide entries have explicit UniProt cautions that
their sequences have not been found in the proteome or mapped to the reference
genome. These exclusions concern the question of complete encoded products;
they do not prove that the reported peptides never existed.

We retained eight entries, including all ties through 31 residues, and sorted
them by residue count with accession as a deterministic tie-breaker. We kept
the 512 screen survivors above 31 residues in the audit file but **did not
manually certify them**. We do not claim to count all bona fide human proteins
under 100 residues.

The retained list includes mitochondrial-derived peptides whose production is
not fully understood. Protein-existence level 1 does not establish an isolated
stable fold, resolve every sequence caveat, or guarantee equivalent experimental
support across entries. Our audit does not include unreviewed records, mature
products embedded in longer precursors, or every short ORF in the literature.
All chart lengths refer to default sequences as listed in the database, without
subtracting initiator methionines or applying other processing annotations.

## Files and reproduction

- `audit.csv`: all 756 records, sequences, evidence levels and screening decisions.
- `shortest.csv`: eight retained short entries and their editorial qualifications.
- `snapshot.json.gz`: frozen full JSON annotations from the two API pages.
- `sources.json`: query, retrieval time, release, URLs and SHA-256 checksums.
- `review.json`: authored decisions for all 14 short screen survivors.
- `summary.json`: generated counts and chart values.
- `analyze.py`: standard-library Python script; no paid tools or dependencies.
- `reproduction.zip`: all eight files above in one downloadable archive.

Download these files into one directory. With Python 3.11 or newer, run:

```sh
python analyze.py --output-dir reproduced
```

The script verifies the snapshot checksum, query population, sequence lengths,
unique accessions, review coverage and reviewed entry versions. It reproduces
the CSV files and summary offline. `python analyze.py --fetch` retrieves the live
database and replaces the snapshot and source metadata; use a separate copy for
this. A new release or changed reviewed entry version requires an updated review.

UniProt is the source of the sequences and annotations and makes its data
available under [CC BY 4.0](https://www.uniprot.org/help/license).
