# How many genes do humans have?

> Humans have 19,442 protein-coding genes in GENCODE Release 50. The broader human reference annotation contains 78,733 gene entries when non-coding genes, pseudogenes, and other categories are included.

Humans have about 20,000 protein-coding genes. GENCODE Release 50 gives a more precise count of 19,442, using its definition of protein-coding genes on the main reference chromosomes. Including non-coding genes, pseudogenes, and other annotation categories brings the same catalog to 78,733 entries.

These numbers answer different questions. A protein-coding gene contains instructions for a protein, while a non-coding RNA gene produces RNA that is not translated into a protein. A genome annotation also records uncertain and gene-like sequences. The total number of annotated entries therefore exceeds the number of established functional genes, and neither count describes how many physical gene copies a person inherits.

## What does the human gene count include?

GENCODE v50 lists 19,442 protein-coding genes and 78,733 gene entries overall. Released in June 2026, it uses the GRCh38.p14 reference assembly. An assembly is the reference DNA sequence; an annotation identifies genes and other features along that sequence.

The statistics cover chromosomes 1–22, X, Y, and MT. They exclude alternate loci and patches, which provide additional representations or corrections of genomic regions. This is a reference catalog, rather than a count of the gene copies in one person's cells. Table 1 accounts for all entries by annotation category; Figure 1 compares the four largest summary categories with the overall total.

  <caption><strong>Table 1. Human gene entries by annotation category in GENCODE v50.</strong> Counts were checked September 19, 2026; summary categories follow GENCODE's published counting rules and sum to the overall total.</caption>
  <thead>
    <tr>
      <th scope="col">Annotation category</th>
      <th scope="col" align="right">GENCODE v50 entries</th>
      <th scope="col">Scope</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Protein-coding genes</td>
      <td align="right">19,442</td>
      <td>Headline count, excluding readthrough genes</td>
    </tr>
    <tr>
      <td>Long non-coding RNA summary</td>
      <td align="right">35,885</td>
      <td>Includes entries awaiting experimental confirmation</td>
    </tr>
    <tr>
      <td>Small non-coding RNA summary</td>
      <td align="right">7,608</td>
      <td>Includes mitochondrial RNA genes and rRNA pseudogenes</td>
    </tr>
    <tr>
      <td>Pseudogene summary</td>
      <td align="right">14,702</td>
      <td>Includes immune-receptor pseudogenes</td>
    </tr>
    <tr>
      <td>Readthrough genes</td>
      <td align="right">665</td>
      <td>Coding annotations spanning neighboring genes</td>
    </tr>
    <tr>
      <td>Coding immune-receptor segments</td>
      <td align="right">412</td>
      <td>Immunoglobulin and T-cell receptor segments</td>
    </tr>
    <tr>
      <td>Artifact entries</td>
      <td align="right">19</td>
      <td>Annotations classified as artifacts</td>
    </tr>
    <tr>
      <td><strong>All annotated entries</strong></td>
      <td align="right"><strong>78,733</strong></td>
      <td><strong>Total across the categories above</strong></td>
    </tr>
  </tbody>

![GENCODE v50 gene entries: 78,733 overall, with the four largest summary categories shown below the total.](/images/charts/human-gene-types-distribution.webp "**Figure 1. Human gene counts in GENCODE v50.** The overall total includes the four categories shown beneath it and the smaller categories listed in Table 1. Source: published GENCODE v50 statistics.")

Readthrough annotations describe transcripts that extend across neighboring gene loci. GENCODE separates these from its headline coding count: 20,107 entries with the `protein_coding` classification minus 665 readthrough entries gives 19,442. Keeping that exclusion explicit prevents a difference in counting rules from being mistaken for a scientific disagreement.

We summed the published counts for the two broad non-coding RNA groups to obtain 43,493 entries combined. This is our calculation from published summary counts, not a separate experimental census. The long-RNA group includes 34,866 `lncRNA` entries and 1,019 `TEC` entries, meaning “to be experimentally confirmed”; the small-RNA group includes 497 `rRNA_pseudogene` entries. The labels therefore do not establish a biological function for every entry. GENCODE's pseudogene summary includes immunoglobulin and T-cell receptor pseudogenes, while rRNA pseudogenes are grouped with small RNAs.

Nor does a count of non-coding RNA genes describe all [non-coding DNA](/guides/human-genome-junk-dna). That broader term includes introns within protein-coding genes and sequences between genes. Gene classification and the fraction of the [genome](/guides/human-genome-size) that encodes proteins are different measurements.

## Why do GENCODE, RefSeq, and HGNC report different counts?

Human protein-coding gene counts differ because databases use different evidence, inclusion rules, and reference sequences. GENCODE and NCBI RefSeq annotate genomes, while the HUGO Gene Nomenclature Committee (HGNC) maintains approved gene names and symbols. Table 2 compares their reported counts alongside the intersection and union from a published three-catalog study.

  <caption><strong>Table 2. Human protein-coding gene counts across databases and a published catalog comparison.</strong> The first three rows come from GENCODE, RefSeq, and the HGNC resource paper, respectively. The final two compare Ensembl/GENCODE, RefSeq, and UniProtKB. These are dated snapshots, not measurements synchronized to one release.</caption>
  <thead>
    <tr>
      <th scope="col">Source and snapshot</th>
      <th scope="col" align="right">Protein-coding count</th>
      <th scope="col">What is counted</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GENCODE v50, June 2026</td>
      <td align="right">19,442</td>
      <td>Main reference chromosomes; readthrough genes excluded</td>
    </tr>
    <tr>
      <td>NCBI RefSeq RS<em>2025</em>08</td>
      <td align="right">19,890</td>
      <td>GRCh38.p14 primary assembly</td>
    </tr>
    <tr>
      <td>HGNC, September 3, 2025</td>
      <td align="right">19,294</td>
      <td>Genes with approved protein-coding symbols</td>
    </tr>
    <tr>
      <td>Maquedano and colleagues, 2025: intersection</td>
      <td align="right">19,268</td>
      <td>Genes classified as coding in all three catalogs studied</td>
    </tr>
    <tr>
      <td>Maquedano and colleagues, 2025: union</td>
      <td align="right">21,871</td>
      <td>Genes classified as coding in at least one catalog studied</td>
    </tr>
  </tbody>

[Download the cross-database human gene count data (CSV)](/data/human-gene-counts-by-database-2026.csv)

Assembly scope alone can change a total within one database. RefSeq's August 2025 annotation reports 19,890 coding genes on the GRCh38 primary assembly, 20,076 across its full reported GRCh38 scope, and 20,070 on T2T-CHM13v2.0. The larger GRCh38 figure includes genes represented on alternate loci or patches; it is not obtained by simply adding the counts for each assembly unit, because some genes occur on multiple units.

The three-catalog study addresses a different question: whether curators agree that a locus encodes a protein. Maquedano and colleagues found 2,603 genes whose coding classification differed among the catalogs. Their analysis used GENCODE v45, so its consensus count is not an updated estimate for v50. The authors also showed that removing readthrough genes and immunoglobulin fragments reduced disagreement.

A precise gene count consequently needs four details: the database, release, reference assembly, and inclusion rule. Without them, a difference of several hundred genes may reflect catalog scope as much as biological evidence.

## Why do some sources say 20,000–25,000 genes?

The familiar 20,000–25,000 range is a historical estimate of human protein-coding genes from the Human Genome Project's finished-sequence analysis in 2004. It does not include every class of non-coding gene. Table 3 places this estimate between earlier expectations and the release-specific GENCODE count.

  <caption><strong>Table 3. Selected historical estimates of the human protein-coding gene count.</strong> Historical figures come from NHGRI's 2001 and 2004 announcements; the modern count comes from GENCODE. The rows compare an expectation, estimated ranges, and an annotation count, rather than repeated measurements using one method.</caption>
  <thead>
    <tr>
      <th scope="col">Period or annotation</th>
      <th scope="col" align="right">Protein-coding estimate or count</th>
      <th scope="col">Basis</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Before the draft genome</td>
      <td align="right">About 100,000</td>
      <td>Widely held expectation recalled in the 2004 NHGRI announcement</td>
    </tr>
    <tr>
      <td>2001 draft genome</td>
      <td align="right">30,000–35,000</td>
      <td>Initial sequence analysis</td>
    </tr>
    <tr>
      <td>2004 finished-sequence analysis</td>
      <td align="right">20,000–25,000</td>
      <td>Revised estimate using a more complete sequence</td>
    </tr>
    <tr>
      <td>GENCODE v50, June 2026</td>
      <td align="right">19,442</td>
      <td>Release-specific annotation count</td>
    </tr>
  </tbody>

The 2004 analysis reported 19,599 confirmed protein-coding genes and another 2,188 predicted coding segments. Better sequence coverage helped resolve errors in earlier gene models, while comparisons with other organisms and improved computational methods strengthened annotation. Subsequent curation has continued to reconsider whether individual loci encode proteins.

A declining estimate does not mean that humans lost thousands of genes over these decades. The change reflects what researchers could identify and support with evidence. Nor is every future release required to contain fewer genes: curators can recognize previously missed loci as well as revise existing classifications.

For general explanations, “about 20,000 protein-coding genes” remains appropriate. MedlinePlus uses about 19,900 while also explaining the Human Genome Project's older range. Exact values are useful when tied to the annotation that produced them.

## How much has the human gene catalog changed since 2024?

We compared GENCODE v45 with v50 and found a net increase of 47 protein-coding entries, from 19,395 to 19,442. By matching gene identifiers, we found more movement than this difference alone suggests: 60 identifiers entered the counted coding set and 13 left it, while 19,382 remained coding at both endpoints. These are results of our comparison of published annotations, not new experimental measurements or a count of newly discovered biological genes.

Table 4 follows the six releases from January 2024 to June 2026. All use GRCh38.p14, and all counts use the comprehensive annotation of the main reference chromosomes. The coding count excludes genes tagged as readthroughs in each release.

  <caption><strong>Table 4. Human gene counts and coding-set changes across GENCODE v45–v50.</strong> We analysed the six comprehensive reference-chromosome GTF files and checked total, coding, and excluded-readthrough counts against each release's published statistics. Entry and exit counts describe identifiers between adjacent rows; they must not be summed to obtain the v45-to-v50 comparison because an identifier can leave and return. Release dates follow GENCODE's release history. <a href="/data/guides/gencode-v45-v50/release-counts.csv">Download the release totals</a> and <a href="/data/guides/gencode-v45-v50/coding-transitions.csv">transition counts</a>.</caption>
  <thead>
    <tr>
      <th scope="col">Release</th>
      <th scope="col">Date</th>
      <th scope="col" align="right">All entries</th>
      <th scope="col" align="right">Coding</th>
      <th scope="col" align="right">Entering</th>
      <th scope="col" align="right">Leaving</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>v45</td>
      <td>January 2024</td>
      <td align="right">63,187</td>
      <td align="right">19,395</td>
      <td align="right">Not compared</td>
      <td align="right">Not compared</td>
    </tr>
    <tr>
      <td>v46</td>
      <td>May 2024</td>
      <td align="right">63,086</td>
      <td align="right">19,411</td>
      <td align="right">17</td>
      <td align="right">1</td>
    </tr>
    <tr>
      <td>v47</td>
      <td>October 2024</td>
      <td align="right">78,724</td>
      <td align="right">19,433</td>
      <td align="right">26</td>
      <td align="right">4</td>
    </tr>
    <tr>
      <td>v48</td>
      <td>May 2025</td>
      <td align="right">78,686</td>
      <td align="right">19,435</td>
      <td align="right">11</td>
      <td align="right">9</td>
    </tr>
    <tr>
      <td>v49</td>
      <td>September 2025</td>
      <td align="right">78,691</td>
      <td align="right">19,433</td>
      <td align="right">1</td>
      <td align="right">3</td>
    </tr>
    <tr>
      <td>v50</td>
      <td>June 2026</td>
      <td align="right">78,733</td>
      <td align="right">19,442</td>
      <td align="right">9</td>
      <td align="right">0</td>
    </tr>
  </tbody>

The coding total stayed close to 19,400 even as the broader catalog expanded. The largest step in Table 4 occurs between v46 and v47: all entries increase by 15,638, while coding entries increase by 22. Over the same interval, entries classified specifically as `lncRNA` rise from 19,258 to 34,914. Changes in an all-gene total can therefore be dominated by RNA annotation rather than protein-coding genes.

For Figure 2, we compiled both counts back to 2014, using the latest published GRCh38 release in each calendar year. The full catalog grows from 60,155 entries in v21 to 78,733 in v50, with a dip in 2016 and a pronounced increase in 2024. The published protein-coding count stays near 20,000, from 19,881 in v21 to 19,442 in v50. The larger total includes non-coding RNA genes, pseudogenes, and other categories.

![GENCODE annual counts from 2014 to 2026: all annotated gene entries rise from 60,155 to 78,733, while published protein-coding counts remain near 20,000, ending at 19,442.](/images/charts/human-gene-catalog-growth-2014-2026.webp "**Figure 2. All annotated genes and protein-coding genes in GENCODE, 2014–2026.** Both lines use the latest published GRCh38 release of each year; 2026 uses v50, released in June and current on September 19. Coding counts follow each release's published definition; the selected releases from 2022 onward exclude readthrough genes. Lines connect annual annotation snapshots, not biological discoveries or losses. We compiled these counts from GENCODE release statistics.")

We began the timeline in 2014 to keep our comparison within the GRCh38 assembly series. Assembly patch versions change, but every selected statistics page counts only the main reference chromosomes. We used the published summary count for the coding series without retrospectively harmonizing definitions: the selected v42–v50 releases exclude readthrough genes, while v21–v39 do not separately subtract them. The change between the 2021 and 2022 snapshots therefore includes a counting-rule change, not simply a loss of coding genes. Annotation rules and evidence also evolve, so these lines describe catalog counts rather than a fixed set of equally validated functional genes. Release months vary, and intermediate releases are omitted. The [downloadable chart data](/data/guides/gencode-gene-count-history-2014-2026.csv) records both counts, the coding readthrough policy, release, month, assembly version, source URL, and source-page checksum.

In Figure 3, we grouped identifiers by the annotation changes we observed when they entered or left the coding set between v45 and v50. Of the 60 entering identifiers, 30 were already present under another biotype: 16 as lncRNA and 14 as pseudogenes. One remained classified as protein-coding but lost its readthrough flag, and 29 were absent from the v45 gene records. In the opposite direction, four identifiers changed biotype, three acquired a readthrough flag, and six were absent from v50.

![Coding-set changes from GENCODE v45 to v50: entries comprise 30 biotype changes, 29 previously absent identifiers, and one removed readthrough flag; exits comprise six absent identifiers, four biotype changes, and three added readthrough flags.](/images/charts/human-coding-gene-turnover-v45-v50.webp "**Figure 3. Observed coding-set changes between GENCODE v45 and v50.** We analysed reference-chromosome gene records, excluding readthrough genes. Flag changes refer to the readthrough annotation. Bars count identifiers entering or leaving the set, not independently confirmed biological discoveries or losses.")

Identifier changes require particular care. We found that 24 of the 29 entering identifiers absent from v45 overlap a same-strand gene span already annotated in v45. All six disappearing identifiers overlap another gene span in v50. Overlap alone does not prove that two records represent the same gene, but it shows why “absent identifier” cannot be equated with “previously unknown gene.”

PAXX provides a concrete example: v47 records `ENSG00000148362`, while v48 records `ENSG00000310560` with the same gene name, strand, and genomic span. Identifier matching records an exit and an entry despite this continuity. CAST shows a related complication: its older identifier changes to lncRNA in v47 while another coding identifier with the same name appears at an overlapping location. That change does not establish that CAST ceased to encode a protein.

We counted only GTF gene rows and matched identifiers after removing their numeric version suffix. We preserved any chromosome-Y pseudoautosomal suffix and did not match by gene name. An identifier coding at both endpoints counts as retained even if its transcripts or coordinates changed. We flagged missing identifiers for same-strand span overlap, including introns, without automatically assigning them to splits, mergers, or replacements. We have not adjudicated the underlying cause of every annotation change.

The [gene-level comparison](/data/guides/gencode-v45-v50/changed-coding-identifiers.csv), [full methods](/data/guides/gencode-v45-v50/methods.txt), [analysis script](/data/guides/gencode-v45-v50/analyze.py), [source manifest](/data/guides/gencode-v45-v50/sources.json), and [run provenance](/data/guides/gencode-v45-v50/provenance.json) provide the records, inclusion rules, download URLs, and checksums needed to reproduce the results. Filter the gene-level file to releases 45 and 50 for the endpoint comparison shown in Figure 3.

## Does a human cell have 20,000 genes or 40,000?

About 20,000 describes distinct protein-coding gene loci, meaning positions in the genome, rather than all their inherited copies. For most nuclear genes, a person inherits one copy from each parent. Different sequence versions of the same gene are called alleles.

Most human body cells are diploid, with 46 chromosomes arranged in 23 pairs. Sperm and egg cells carry a single chromosome set. Two chromosome sets provide two copies of most genes without doubling the number of distinct gene identities. A gene count is therefore different from a chromosome count: each chromosome contains many genes.

The reference annotation also includes both X and Y and the mitochondrial chromosome, MT. It should not be multiplied by two to obtain an exact gene-copy count for every cell. The guides to [genes per chromosome](/guides/how-many-genes-are-in-a-chromosome) and [mitochondrial genes](/guides/how-many-genes-in-mitochondrial-dna) explain these distinct parts of the genome.

## Do humans have more genes than other animals?

Humans do not have an exceptionally large protein-coding gene set. Several animal reference annotations contain more coding genes, while others contain fewer. Table 5 retains the Ensembl release snapshots used for this guide rather than treating gene counts as permanent properties of each species.

  <caption><strong>Table 5. Protein-coding gene counts in selected species.</strong> The human row uses GENCODE. Other rows use the named Ensembl and Ensembl Plants annotations, recorded August 17, 2026.</caption>
  <thead>
    <tr>
      <th scope="col">Organism</th>
      <th scope="col" align="right">Coding genes</th>
      <th scope="col">Reference annotation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Human</td>
      <td align="right">19,442</td>
      <td>GENCODE v50, GRCh38.p14</td>
    </tr>
    <tr>
      <td>Mouse</td>
      <td align="right">22,081</td>
      <td>Ensembl 116, GRCm39</td>
    </tr>
    <tr>
      <td>Dog</td>
      <td align="right">20,567</td>
      <td>Ensembl 116, ROS<em>Cfam</em>1.0</td>
    </tr>
    <tr>
      <td>Domestic cat</td>
      <td align="right">19,209</td>
      <td>Ensembl 116, F.catus<em>Fca126</em>mat1.0</td>
    </tr>
    <tr>
      <td>Cattle</td>
      <td align="right">20,848</td>
      <td>Ensembl 116, ARS-UCD2.0</td>
    </tr>
    <tr>
      <td>Zebrafish</td>
      <td align="right">25,592</td>
      <td>Ensembl 116, GRCz11</td>
    </tr>
    <tr>
      <td>Fruit fly</td>
      <td align="right">13,986</td>
      <td>Ensembl 116, BDGP6.46</td>
    </tr>
    <tr>
      <td><em>Arabidopsis thaliana</em></td>
      <td align="right">27,655</td>
      <td>Ensembl Plants 63, TAIR10/Araport11</td>
    </tr>
    <tr>
      <td>Baker's yeast</td>
      <td align="right">6,600</td>
      <td>Ensembl 116, R64-1-1</td>
    </tr>
  </tbody>

Annotation methods and evidence differ among species, so the counts in Table 5 compare the scale of cataloged gene sets rather than provide a standardized measure of biological complexity.

Gene number also differs from the number of gene products. A gene can give rise to several RNA transcripts, and coding transcripts can specify different protein sequences. GENCODE v50 reports 278,455 protein-coding transcripts and 172,117 distinct translations. The translation statistic counts distinct sequences within each gene and sums across genes; it is not a measurement of the proteins present in a particular cell.

The [human transcriptome](/guides/human-transcriptome-size) and [protein-count](/guides/number-of-proteins) guides develop these distinctions. Counting genes describes one level of genome organization; understanding their products also requires knowing which transcripts are expressed, where they are expressed, and how their proteins are processed.
