How many transcripts are in the human transcriptome?

Matic BrozComputational chemist
The current human reference annotation contains 644,292 transcript models on the main chromosomes. This is a catalog of RNA isoforms that could be produced, not a count of the transcripts expressed in one cell.
A cell expresses a subset of the catalog and may contain many physical copies of the same RNA. The human transcriptome size therefore depends on whether the unit is an annotated isoform, an expressed isoform, or an individual RNA molecule.
How many transcripts are in the human transcriptome?
GENCODE Release 50 annotates 644,292 human transcript models on the main chromosomes.[1]
The same release contains 78,733 annotated genes. Its transcript total includes 278,455 protein-coding transcripts, 191,063 long non-coding RNA transcripts, 91,818 transcripts expected to undergo nonsense-mediated decay, and other RNA and pseudogene categories.[1] A transcript model is one annotated RNA isoform, so several models can come from the same gene.
GENCODE also reports 172,117 distinct translations. Multiple protein-coding transcript models can resolve to the same translated sequence, so this figure is lower than the 278,455 coding-transcript count.[1] It is a protein-sequence count, not another estimate of the number of RNAs.
All five bars use the same GENCODE release and main-chromosome scope. They show why gene, transcript, and translation totals answer different questions rather than competing estimates of one quantity.[1]
Why do transcript databases report different totals?
GENCODE, Ensembl, RefSeq, and RNAcentral report different totals because they cover different sequence sets and apply different annotation rules.
GENCODE Release 50 corresponds to Ensembl Release 116, and GENCODE is the default human gene set displayed by Ensembl.[2][3] Their current human transcript statistics should not be presented as two independent censuses.
The latest RefSeq report for GRCh38.p14 lists 186,355 mRNA and non-coding RNA models on the primary assembly, excluding pseudogene transcripts. It separately lists 1,828 pseudogene transcripts.[4] This total is not directly comparable with GENCODE's 644,292 because RefSeq and GENCODE differ in their evidence pipelines, categories, and treatment of alternative models.
RNAcentral Release 26 groups 600,225 human non-coding RNA transcripts into 103,814 predicted genes across 56 RNA types.[5] It aggregates non-coding RNA sequences from specialist resources, so its 600,225 figure is neither a whole-transcriptome total nor an alternative count of GENCODE's transcript models.
A transcript count by species needs the same annotation pipeline, release, assembly scope, and inclusion rules for every organism. Without those controls, a better-studied species can appear to have a larger transcriptome because it has been annotated more deeply. This guide therefore does not rank species by transcript count.
How many transcripts are expressed in one human cell?
No single human cell expresses all 644,292 annotated transcript models, and there is no universal per-cell isoform count.
A transcriptome is the collection of gene readouts present in a particular cell or sample. Cell types express different sets of genes, and the number detected changes with sequencing depth, RNA selection, capture efficiency, and the threshold used to call expression.[6] Salmon estimates transcript abundance from RNA sequencing reads, but its output is still specific to the sample and reference annotation supplied.
Distinct expressed isoforms are also different from physical molecules. A calibrated study of human U-2 OS cells estimated about 300,000 mRNA molecules per cell, with individual mRNAs ranging from fewer than one to about 3,500 copies per cell on average.[7] Those 300,000 molecules include repeated copies of the same transcript and exclude abundant non-messenger RNAs such as rRNA and tRNA. The broader RNA composition varies with cell size, type, and activity.
Are transcripts, isoforms, translations, and RNA molecules the same count?
A transcript isoform is an RNA sequence model, a translation is the protein sequence encoded by a coding transcript, and an RNA molecule is one physical copy inside a biological sample.
One gene can have several transcript isoforms through alternative promoters, splicing, or polyadenylation. Two isoforms may encode different proteins, the same protein, or no protein. This is why the current annotation contains 278,455 protein-coding transcript models but only 172,117 distinct translations.[1] The broader protein count adds further levels, including proteoforms and repeated protein molecules.
Physical RNA copies also turn over continuously. Their abundance can change while the reference annotation stays fixed, and their half-lives range from minutes to hours. Reference transcriptome statistics describe the catalog; expression measurements describe a sample at a particular time.
Sources▼
- Human release statistics (version 50) GENCODE · August 10, 2026. https://www.gencodegenes.org/human/stats.html
- The Human GENCODE Gene Set release history GENCODE · August 10, 2026. https://www.gencodegenes.org/human/releases.html
- How to access GENCODE data GENCODE · August 10, 2026. https://www.gencodegenes.org/pages/data_access.html
- Homo sapiens Annotation Release GCF_000001405.40-RS_2025_08 NCBI RefSeq · 2025. https://www.ncbi.nlm.nih.gov/refseq/annotation_euk/Homo_sapiens/GCF_000001405.40-RS_2025_08/
- RNAcentral Release 26 RNAcentral · 2025. https://blog.rnacentral.org/2025/10/rnacentral-release-26.html
- Transcriptome Fact Sheet National Human Genome Research Institute · August 10, 2026. https://www.genome.gov/about-genomics/fact-sheets/Transcriptome-Fact-Sheet
- The Stress Granule Transcriptome Reveals Principles of mRNA Accumulation in Stress Granules Molecular Cell · 2017. https://pmc.ncbi.nlm.nih.gov/articles/PMC5728175/

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.