# Introns vs exons: what is the difference?

> Exons are the parts of a gene kept in the mature RNA; introns are removed by splicing. Learn how splicing works, how much of a human gene is intron, why exons are not the same as coding DNA, and what happens when splicing fails.

## What is the difference between introns and exons?

Exons are the segments of a gene that remain in the mature RNA. Introns are the segments between them that are copied into the first RNA transcript and then cut out by splicing. After the introns are removed, the exons are joined end to end, and in a protein-coding gene that joined RNA becomes the messenger RNA (mRNA) that ribosomes translate.

A common shortcut says exons code for protein and introns do not. That is only half right. Introns are never translated, but exons are not all coding either: the first and last exons usually carry untranslated regions (UTRs) that stay in the mRNA without being read into amino acids. In reviewed human protein-coding transcripts, only about 44% of distinct exon sequence is protein-coding.

| Feature | Exon | Intron |
| --- | --- | --- |
| Present in the gene (DNA) | Yes | Yes |
| Present in the first RNA copy (pre-mRNA) | Yes | Yes |
| Present in the mature mRNA | Yes | No, removed by splicing |
| Translated into protein | Coding parts only; UTRs are not | Never |
| Typical human length | Median 131 bp | Median 1,747 bp |
| Number per human mRNA | Median 9 | Median 8 (one fewer than exons) |
| Share of a human gene's exon plus intron sequence | About 5% | About 95% |
| Boundary signals | Defined by the splice sites of the flanking introns | Usually starts with GT and ends with AG in DNA |

The lengths and counts come from 19,116 human protein-coding genes with reviewed or validated NCBI RefSeq records. The share of exon and intron sequence is our calculation from the same dataset, explained [below](#how-much-of-a-human-gene-is-intron).

![A four-exon gene is transcribed into pre-mRNA; splicing removes three introns and joins the exons, and only the coding sequence is translated.](/images/guides/introns-vs-exons/gene-to-mrna-exons-introns.webp "**Figure 1. From gene to mRNA to protein.** Exons stay in mature mRNA; introns are transcribed and then spliced out. The untranslated regions stay in the RNA but are not translated. Schematic, not to scale.")

## What is an exon?

An exon is any segment of a gene that survives splicing and ends up in the mature RNA. The name comes from Walter Gilbert, who proposed "exon" for expressed regions and "intron" for intragenic regions in a 1978 essay titled *Why genes in pieces?*. Exons exist in non-coding RNA genes too, but most discussions refer to [protein-coding genes](/guides/how-many-genes-do-humans-have).

Human exons are short. Half of all exons in reviewed human mRNAs are 131 bp or shorter, and the mean for exons other than the last is 159 bp. The last exon is usually the longest because it often carries a long 3' UTR; this pulls the mean for all exons up to 311 bp. The longest exon in the dataset is exon 13 of GRIN2B at 27,303 bp, and the transcript with the most exons is from TTN, the gene for [titin](/guides/largest-protein), with 363.

Protein-coding sequence is only part of the exon picture. Distinct exon sequence in reviewed human mRNAs totals 59.3 million bp, of which 25.8 million bp is coding. The rest is UTR sequence, which influences how stable an mRNA is, where it goes in the cell, and how efficiently it is translated. That is one reason [mRNA lifetimes](/guides/how-long-does-mrna-last-in-a-cell) can differ widely between genes with similar proteins.

## What is an intron?

An intron is a segment of a gene that is transcribed into RNA and then removed before the RNA is used. Introns were discovered in 1977, when the laboratories of Phillip Sharp and Richard Roberts independently found that adenovirus mRNAs were stitched together from sequences lying far apart on the viral genome. Sharp and Roberts shared the 1993 Nobel Prize in Physiology or Medicine for the discovery of split genes.

Most introns in human genes are recognized by short sequence signals rather than by their overall content. In DNA, an intron almost always begins with GT and ends with AG (GU and AG in the RNA). An analysis of verified mammalian splice sites estimated that about 99.2% of splice site pairs are GT-AG, 0.7% are GC-AG, and 0.05% are AT-AC. Other signals sit inside the intron: a branch point containing an adenosine, usually a few dozen bases upstream of the 3' end, and a stretch rich in pyrimidines (C and U) between the branch point and the final AG.

Human introns vary enormously in length. The median is 1,747 bp, but the mean is 6,938 bp because a minority of introns are very long; intron 2 of ROBO2 is 1,160,411 bp. No human intron removed by the normal splicing machinery in the dataset is shorter than 30 bp. The 26-nucleotide intron in XBP1 is shorter, but it is cut out by the enzyme IRE1 during cellular stress rather than by the spliceosome.

A small class of introns follows different rules. U12-type, or minor, introns account for less than 0.5% of introns in any given genome and are removed by a separate minor spliceosome.

## How splicing removes introns

Splicing is carried out by the spliceosome, a large complex made of five small nuclear RNA-protein particles (snRNPs, named U1, U2, U4, U5 and U6) and many additional proteins. The spliceosome assembles on each intron, recognizes the 5' splice site, branch point and 3' splice site, and removes the intron in two chemical steps:

1. The 2' hydroxyl of the branch point adenosine attacks the 5' splice site. This cuts the RNA at the start of the intron and joins the intron's first nucleotide to the branch point, forming a loop called a lariat.
2. The freed end of the upstream exon attacks the 3' splice site. This joins the two exons and releases the intron as a lariat, which is then opened and degraded.

In human cells, splicing mostly happens while the RNA is still being transcribed by RNA polymerase II, so the rate of transcription can influence which splice sites are chosen. Splicing has to be exact to the nucleotide. If a splice site shifts by one or two bases in a coding region, the downstream reading frame changes, and the mRNA usually picks up a premature stop codon and is destroyed by nonsense-mediated decay.

![The branch point adenosine forms an intron lariat in the first splicing step; the free end of exon 1 joins exon 2 in the second step, releasing the lariat.](/images/guides/introns-vs-exons/splicing-mechanism-lariat.webp "**Figure 2. The two steps of splicing.** The intron remains attached to exon 2 after the first reaction. The second reaction joins the flanking exons and releases the intron lariat. U1 and U2 illustrate early splice-site recognition; the spliceosome changes composition during assembly and catalysis. Schematic, not to scale.")

### Worked example: why a genomic sequence gives the wrong protein

A toy gene shows why the exon-intron structure matters when you work with sequences. The two exons below encode a short peptide. The intron between them starts with GT, ends with AG, and contains a branch-point-like sequence and a pyrimidine-rich stretch.
```text
Exon 1   ATGGCTTCAAAG
Intron   GTAAGTATGATTTACTAACTTCTTTTCCTTGCAG
Exon 2   GAATTCTGGTGA
```
If you translate the exons joined together, the result is the intended peptide, `MASKEFW` followed by a stop codon. If you translate the unspliced sequence, the ribosome's reading frame runs into the intron and reaches a TAA stop codon after nine amino acids, giving `MASKVSMIY`. Real genes behave the same way, only at a much larger scale: a median human gene is about 26,000 bp long, while a median human mRNA is about 2,900 bp.

You can reproduce this in the [DNA to protein translator](/app/dna-to-protein) by pasting the unspliced and spliced sequences in turn. An [ORF finder](/app/orf-finder) run on genomic DNA from a eukaryote will likewise report fragments of the real open reading frame, broken wherever an intron interrupts it.

## How much of a human gene is intron?

Introns make up the great majority of a typical human protein-coding gene. Across reviewed human transcripts, distinct exon sequence totals 59,281,518 bp and distinct intron sequence totals 1,095,434,245 bp. Dividing the first number by their sum gives:

59,281,518 / (59,281,518 + 1,095,434,245) = 5.1%

So about 5% of the exon plus intron sequence of human protein-coding genes is exon, and about 95% is intron. Put another way, there is roughly 18.5 times as much intron sequence as exon sequence. Only 25.8 million bp of the exon total is protein-coding. That is why protein-coding sequence covers only about 1% of the [3.1 billion bp human genome](/guides/human-genome-size); ENCODE put the figure at 1.22%.

| Measure (reviewed human RefSeq, 2019) | Exons | Introns |
| --- | --- | --- |
| Distinct features | 159,652 | 148,092 |
| Median per transcript | 9 | 8 |
| Median length | 131 bp | 1,747 bp |
| Mean length | 311 bp | 6,938 bp |
| Shortest | 2 bp | 30 bp (spliceosomal) |
| Longest | 27,303 bp (GRIN2B exon 13) | 1,160,411 bp (ROBO2 intron 2) |
| Total distinct sequence | 59.3 million bp | 1,095.4 million bp |

The dataset covers 19,116 genes and 49,632 mRNAs with reviewed or validated RefSeq status as of January 2019. "Distinct" counts each exon or intron once even when it appears in several transcript isoforms. Newer annotations add transcripts but, according to the authors, typical exon and intron lengths had already stabilized.

These proportions explain why some genes take hours to transcribe. The [DMD gene](/guides/largest-and-smallest-human-genes) has 79 exons spread over at least 2.3 million bp, and RNA polymerase needs an estimated 16 hours to copy it, most of that time spent on introns. Splicing of its 5' end begins before transcription is finished. At the scale of the whole genome, ENCODE estimated that introns of protein-coding genes cover about 37% of human DNA, which makes the [junk DNA](/guides/human-genome-junk-dna) debate partly a debate about introns.

![Donut chart of combined exon and intron sequence: introns 94.9%, untranslated exon sequence 2.9%, and protein-coding exon sequence 2.2%.](/images/charts/human-gene-exon-intron-proportion.webp "**Figure 3. Exons occupy about 5% of the combined exon and intron sequence.** Slice sizes use the exact non-redundant totals from reviewed or validated human RefSeq protein-coding transcripts in January 2019. This is an aggregate across the dataset, not the structure of a single gene. UTR sequence is calculated as total exon sequence minus coding sequence. Source: [Piovesan et al. (2019)](https://pmc.ncbi.nlm.nih.gov/articles/PMC6549324/), Table 2.")

![Human exon and intron lengths on a shared linear scale: medians of 131 and 1,747 bp, and means of 311 and 6,938 bp.](/images/charts/human-exon-intron-lengths.webp "**Figure 4. Introns are much longer than exons.** Medians and means count all exon and intron entries across the reviewed or validated transcript dataset, rather than collapsing repeated features across isoforms. Source: [Piovesan et al. (2019)](https://pmc.ncbi.nlm.nih.gov/articles/PMC6549324/), Table 2.")

## Alternative splicing: one gene, many exon combinations

Because exons are separate pieces, a cell can join them in more than one way. This is alternative splicing, and it is the norm in human genes rather than an exception. RNA sequencing studies published in 2008 estimated that 92% to 94% of human genes, or about 95% of multi-exon genes, undergo alternative splicing. Most of these events differ between tissues.

The main patterns are:

- Exon skipping, where a cassette exon is included in some transcripts and left out of others
- Mutually exclusive exons, where one of two neighboring exons is used but never both
- Alternative 5' or 3' splice sites, which lengthen or shorten an exon from one end
- Intron retention, where an intron stays in the mature RNA

Alternative splicing is a major reason one gene can produce several proteins. It is also why counts differ depending on what you count: there are about 20,000 protein-coding genes but many more annotated transcripts in the [human transcriptome](/guides/human-transcriptome-size) and translated sequences in the [human proteome](/guides/number-of-proteins). Not every alternative transcript makes a distinct, functional protein, so isoform counts are an upper bound on protein diversity rather than a measure of it.

![Four alternative splicing patterns: exon skipping, mutually exclusive exons, alternative splice sites, and intron retention, each yielding different RNA isoforms.](/images/guides/introns-vs-exons/alternative-splicing-patterns.webp "**Figure 5. Four patterns of alternative splicing.** The alternative-site panel shows two possible 5′ splice sites; 3′ splice sites can also vary. Intron retention can introduce a premature stop codon. Different RNA isoforms do not necessarily produce distinct functional proteins. Schematic, not to scale.")

## Why do introns exist?

Nobody has a single, settled answer. Gilbert's 1978 essay argued that genes built from separate exons can evolve faster, because recombination within long introns can shuffle whole exons, and the domains they encode, between genes without breaking the coding sequence. Whether most introns were present in early life and lost from bacteria, or were inserted later into eukaryotic genes, is still debated.

Whatever their origin, introns now do several jobs. Their sequences contain regulatory elements, and in many eukaryotes an intron can raise gene expression by increasing transcription, nuclear export, mRNA stability or translation, an effect called intron-mediated enhancement. Introns also make alternative splicing possible. The same split structure carries a cost: every intron is another place where a mutation can derail splicing, which is why splicing errors are a recurring cause of inherited disease.

Intron density varies widely across life. Bacteria have few introns, and the ones they have are self-splicing group I and group II introns rather than spliceosomal introns. Budding yeast, a eukaryote with a compact genome, had only 228 introns in a curated 1999 catalog. Human protein-coding transcripts contain more than 148,000 distinct introns.

## What happens when splicing goes wrong

Because the spliceosome reads short signals, a single base change can alter which exons end up in the mRNA. Mutations can destroy a splice site, create a new one, or disrupt the enhancer and silencer sequences that help the spliceosome choose between sites. One of the first known examples was a mutation that creates an alternative 3' splice site in the β-globin gene and causes beta-thalassemia.

Spinal muscular atrophy (SMA) shows how a change that does not alter the protein sequence can still cause disease. The SMN1 and SMN2 genes encode the same protein, but a C-to-T change in SMN2 exon 7 is translationally silent and weakens an exonic splicing enhancer, so most SMN2 transcripts skip exon 7 and make an unstable protein. People with SMA lack working SMN1, and SMN2 cannot fully compensate. Nusinersen, an antisense oligonucleotide that blocks a splicing silencer in SMN2 intron 7 and promotes exon 7 inclusion, was tested in infants with SMA. At the interim analysis of its phase 3 trial, 41% of infants receiving the drug (21 of 51) had a motor-milestone response, compared with none of 27 in the control group.

Duchenne muscular dystrophy runs the logic in reverse. Deletions in DMD often shift the reading frame, and an antisense drug that makes the spliceosome skip one more exon, such as exon 51, can restore the frame and produce a shorter dystrophin that keeps partial function.

## Introns and exons in sequence analysis

The exon-intron structure changes which sequence you should use. A genomic sequence contains introns; an mRNA or cDNA sequence does not. If the goal is to study the encoded protein, predict its structure with a tool such as [ESMFold](/app/esmfold), or design a construct for expression in bacteria, which cannot splice human introns, start from the spliced coding sequence or the protein sequence of the specific isoform you care about.

Comparing an mRNA with its gene by [pairwise sequence alignment](/use-cases/pairwise-sequence-alignment) shows the structure directly: exons align as blocks, separated by long gaps where the introns were. In quantitative PCR, placing a primer across an exon-exon junction is a common way to amplify cDNA but not contaminating genomic DNA, because the junction sequence only exists after splicing. [Primer3](/app/primer3) can design primers around a chosen target region, as covered in the [Primer3 walkthrough](/guides/how-to-use-primer3-online).

For quick checks on a sequence, ProteinIQ has browser tools that cover the steps in this guide:

- [DNA to RNA](/app/dna-to-rna) converts a DNA sequence to its RNA form, turning GT-AG intron ends into GU-AG
- [DNA to protein](/app/dna-to-protein) translates a DNA or mRNA sequence into amino acids
- [ORF finder](/app/orf-finder) locates open reading frames; in eukaryotic genomic DNA, expect them broken into pieces by introns
- [Reverse complement](/app/reverse-complement) flips a gene annotated on the minus strand into its transcribed orientation
