# What are the largest and smallest human genes?

> RBFOX1 is the largest human protein-coding gene by GENCODE v50 genomic span. MLDHR is the shortest under an HGNC-approved protein-coding filter, at 96 base pairs.

By genomic span, **RBFOX1 is the largest annotated human protein-coding gene** in GENCODE v50, at 2,473,539 base pairs. DMD, the textbook answer, ranks fourth at 2,241,933 base pairs.

The smallest gene has no equally simple answer. MLDHR is 96 base pairs under an HGNC-approved protein-coding filter, while the 8-base-pair TRDD1 immune-receptor segment is the shortest gene feature in the unfiltered annotation.

## What is the largest gene in the human genome?

RBFOX1 is the largest human protein-coding gene by genomic span in GENCODE v50, covering 2,473,539 base pairs on chromosome 16.

GENCODE v50 was released in June 2026 for the GRCh38.p14 [human genome](/guides/human-genome-size) and is paired with Ensembl release 116. Its four longest protein-coding gene spans are RBFOX1 at 2,473,539 bp, CNTNAP2 at 2,304,997 bp, PTPRD at 2,298,757 bp, and DMD at 2,241,933 bp. The current Ensembl record uses the same RBFOX1 boundaries.

![Genomic span comparison for RBFOX1, CNTNAP2, PTPRD, DMD, GAGE12B, and MLDHR on a logarithmic scale](/images/charts/human-gene-size-span-comparison.webp)

The chart compares genomic spans, including introns, on a logarithmic scale. RBFOX1 spans about 25,800 times as many bases as MLDHR under the filters described below. We calculated both numbers from the same GENCODE release.

## Is DMD the largest human gene?

No. DMD is exceptionally long, but it ranks fourth among protein-coding genes by genomic span in GENCODE v50.

DMD became the standard answer because a 1995 study described it as the largest gene then known, with 79 exons across at least 2.3 million bases. The study also found that one DMD transcript takes about 16 hours to make. Current annotation catalogs include longer boundaries for RBFOX1, CNTNAP2, and PTPRD.

Even on the same GRCh38.p14 assembly, annotation providers draw some boundaries differently. NCBI RefSeq currently spans RBFOX1 across 2,473,620 bp and DMD across 2,220,167 bp, compared with 2,473,539 bp and 2,241,933 bp in GENCODE. A precise claim therefore needs both the annotation source and release.

Gene length also does not determine protein length. DMD encodes dystrophin, but TTN encodes [titin](/guides/largest-protein), the largest human protein. Introns make up most of the DMD gene span and are removed from the [mature RNA](/guides/human-transcriptome-size).

## What is the smallest gene in humans?

There is no filter-free single answer: TRDD1 is the shortest annotated gene feature at 8 bp, while MLDHR is the shortest HGNC-approved protein-coding gene at 96 bp in GENCODE v50.

TRDD1 is a T-cell receptor diversity segment, not a conventional standalone protein-coding gene. If every row labeled `protein_coding` is accepted, GENCODE also contains an unnamed 27-bp locus, ENSG00000310590. Requiring an HGNC-approved symbol and a protein-coding biotype removes both ambiguities.

Under that rule, MLDHR is one exon and 96 bp long on chromosome 10. NCBI lists it as a validated protein-coding gene, and HGNC approved the symbol in 2024. MLDHR is also known as MP31 because it encodes a 31-amino-acid micropeptide involved in mitochondrial lactate metabolism. Protein length is a separate measurement, covered in our guide to the [smallest proteins](/guides/smallest-protein).

## How was gene length calculated?

Gene length in this guide means the inclusive genomic span of a GENCODE gene feature: end coordinate minus start coordinate plus one.

We downloaded the chromosome annotation for GENCODE v50 on August 10, 2026. For the largest-gene ranking, we kept protein-coding genes and excluded readthrough loci, matching GENCODE's reported set of 19,442 protein-coding genes. For the shortest conventional protein-coding result, we also required an HGNC identifier.

| Measurement       | What it counts                                                     | Why the ranking changes                            |
| ----------------- | ------------------------------------------------------------------ | -------------------------------------------------- |
| Genomic span      | Every base from the annotated gene start to end, including introns | Boundaries vary by assembly and annotation release |
| Transcript length | Exons joined in one RNA isoform                                    | One gene can have many transcripts                 |
| Coding sequence   | Bases translated into protein                                      | UTRs and introns are excluded                      |
| Protein length    | Amino acids in one protein product                                 | This measures the product, not the gene            |

The distinction matters most for long, intron-rich genes. A mature transcript can be far shorter than the genomic region from which it was transcribed. An [ORF finder](/app/orf-finder) can identify candidate coding regions in a DNA sequence; the full genomic span still comes from the gene annotation.
