ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Genetics

How many genes do humans have?

September 19, 2026·Matic Broz, PhD
Conceptual illustration tracing a human body to a chromosome and the exon-intron structure of a gene.

TL;DR

  • As of September 19, 2026, GENCODE v50 lists 19,442 protein-coding genes, excluding readthrough genes from that headline count.
  • The same release contains 78,733 annotated gene entries across all categories; that is not a count of confirmed functional genes.
  • A gene count describes distinct gene loci, not the copies inherited from each parent or the number of proteins a cell makes.
  • Our comparison of GENCODE v45 and v50 finds 60 identifiers entering and 13 leaving the counted coding set, a net increase of 47. These are annotation changes, not 60 confirmed gene discoveries.

Humans have about 20,000 protein-coding genes. GENCODE Release 50 gives a more precise count of 19,442, using its definition of protein-coding genes on the main reference chromosomes. Including non-coding genes, pseudogenes, and other annotation categories brings the same catalog to 78,733 entries.

These numbers answer different questions. A protein-coding gene contains instructions for a protein, while a non-coding RNA gene produces RNA that is not translated into a protein. A genome annotation also records uncertain and gene-like sequences. The total number of annotated entries therefore exceeds the number of established functional genes, and neither count describes how many physical gene copies a person inherits.

What does the human gene count include?

GENCODE v50 lists 19,442 protein-coding genes and 78,733 gene entries overall. Released in June 2026, it uses the GRCh38.p14 reference assembly.[1][2] An assembly is the reference DNA sequence; an annotation identifies genes and other features along that sequence.

The statistics cover chromosomes 1–22, X, Y, and MT. They exclude alternate loci and patches, which provide additional representations or corrections of genomic regions. This is a reference catalog, rather than a count of the gene copies in one person's cells.[3] Table 1 accounts for all entries by annotation category; Figure 1 compares the four largest summary categories with the overall total.

Annotation categoryGENCODE v50 entriesScope
Protein-coding genes19,442Headline count, excluding readthrough genes
Long non-coding RNA summary35,885Includes entries awaiting experimental confirmation
Small non-coding RNA summary7,608Includes mitochondrial RNA genes and rRNA pseudogenes
Pseudogene summary14,702Includes immune-receptor pseudogenes
Readthrough genes665Coding annotations spanning neighboring genes
Coding immune-receptor segments412Immunoglobulin and T-cell receptor segments
Artifact entries19Annotations classified as artifacts
All annotated entries78,733Total across the categories above
Table 1. Human gene entries by annotation category in GENCODE v50. Counts were checked September 19, 2026; summary categories follow GENCODE's published counting rules and sum to the overall total.[1][3]
Figure 1. Human gene counts in GENCODE v50. The overall total includes the four categories shown beneath it and the smaller categories listed in Table 1. Source: published GENCODE v50 statistics. Reuse under CC BY 4.0.

Readthrough annotations describe transcripts that extend across neighboring gene loci. GENCODE separates these from its headline coding count: 20,107 entries with the protein_coding classification minus 665 readthrough entries gives 19,442.[1][3] Keeping that exclusion explicit prevents a difference in counting rules from being mistaken for a scientific disagreement.

We summed the published counts for the two broad non-coding RNA groups to obtain 43,493 entries combined. This is our calculation from published summary counts, not a separate experimental census. The long-RNA group includes 34,866 lncRNA entries and 1,019 TEC entries, meaning “to be experimentally confirmed”; the small-RNA group includes 497 rRNA_pseudogene entries.[1][6] The labels therefore do not establish a biological function for every entry. GENCODE's pseudogene summary includes immunoglobulin and T-cell receptor pseudogenes, while rRNA pseudogenes are grouped with small RNAs.[3]

Nor does a count of non-coding RNA genes describe all non-coding DNA. That broader term includes introns within protein-coding genes and sequences between genes. Gene classification and the fraction of the genome that encodes proteins are different measurements.

Why do GENCODE, RefSeq, and HGNC report different counts?

Human protein-coding gene counts differ because databases use different evidence, inclusion rules, and reference sequences. GENCODE and NCBI RefSeq annotate genomes, while the HUGO Gene Nomenclature Committee (HGNC) maintains approved gene names and symbols. Table 2 compares their reported counts alongside the intersection and union from a published three-catalog study.

Source and snapshotProtein-coding countWhat is counted
GENCODE v50, June 202619,442Main reference chromosomes; readthrough genes excluded
NCBI RefSeq RS20250819,890GRCh38.p14 primary assembly
HGNC, September 3, 202519,294Genes with approved protein-coding symbols
Maquedano and colleagues, 2025: intersection19,268Genes classified as coding in all three catalogs studied
Maquedano and colleagues, 2025: union21,871Genes classified as coding in at least one catalog studied
Table 2. Human protein-coding gene counts across databases and a published catalog comparison. The first three rows come from GENCODE, RefSeq, and the HGNC resource paper, respectively. The final two compare Ensembl/GENCODE, RefSeq, and UniProtKB. These are dated snapshots, not measurements synchronized to one release.[1][7][8][9]

Download the cross-database human gene count data (CSV)

Assembly scope alone can change a total within one database. RefSeq's August 2025 annotation reports 19,890 coding genes on the GRCh38 primary assembly, 20,076 across its full reported GRCh38 scope, and 20,070 on T2T-CHM13v2.0. The larger GRCh38 figure includes genes represented on alternate loci or patches; it is not obtained by simply adding the counts for each assembly unit, because some genes occur on multiple units.[7]

The three-catalog study addresses a different question: whether curators agree that a locus encodes a protein. Maquedano and colleagues found 2,603 genes whose coding classification differed among the catalogs. Their analysis used GENCODE v45, so its consensus count is not an updated estimate for v50. The authors also showed that removing readthrough genes and immunoglobulin fragments reduced disagreement.[9]

A precise gene count consequently needs four details: the database, release, reference assembly, and inclusion rule. Without them, a difference of several hundred genes may reflect catalog scope as much as biological evidence.

Why do some sources say 20,000–25,000 genes?

The familiar 20,000–25,000 range is a historical estimate of human protein-coding genes from the Human Genome Project's finished-sequence analysis in 2004. It does not include every class of non-coding gene.[10] Table 3 places this estimate between earlier expectations and the release-specific GENCODE count.

Period or annotationProtein-coding estimate or countBasis
Before the draft genomeAbout 100,000Widely held expectation recalled in the 2004 NHGRI announcement
2001 draft genome30,000–35,000Initial sequence analysis
2004 finished-sequence analysis20,000–25,000Revised estimate using a more complete sequence
GENCODE v50, June 202619,442Release-specific annotation count
Table 3. Selected historical estimates of the human protein-coding gene count. Historical figures come from NHGRI's 2001 and 2004 announcements; the modern count comes from GENCODE. The rows compare an expectation, estimated ranges, and an annotation count, rather than repeated measurements using one method.[11][10][1][2]

The 2004 analysis reported 19,599 confirmed protein-coding genes and another 2,188 predicted coding segments. Better sequence coverage helped resolve errors in earlier gene models, while comparisons with other organisms and improved computational methods strengthened annotation.[10] Subsequent curation has continued to reconsider whether individual loci encode proteins.[9]

A declining estimate does not mean that humans lost thousands of genes over these decades. The change reflects what researchers could identify and support with evidence. Nor is every future release required to contain fewer genes: curators can recognize previously missed loci as well as revise existing classifications.

For general explanations, “about 20,000 protein-coding genes” remains appropriate. MedlinePlus uses about 19,900 while also explaining the Human Genome Project's older range.[4] Exact values are useful when tied to the annotation that produced them.

How much has the human gene catalog changed since 2024?

We compared GENCODE v45 with v50 and found a net increase of 47 protein-coding entries, from 19,395 to 19,442. By matching gene identifiers, we found more movement than this difference alone suggests: 60 identifiers entered the counted coding set and 13 left it, while 19,382 remained coding at both endpoints. These are results of our comparison of published annotations, not new experimental measurements or a count of newly discovered biological genes.[20][21]

Table 4 follows the six releases from January 2024 to June 2026. All use GRCh38.p14, and all counts use the comprehensive annotation of the main reference chromosomes. The coding count excludes genes tagged as readthroughs in each release.[2][3][22]

ReleaseDateAll entriesCodingEnteringLeaving
v45January 202463,18719,395Not comparedNot compared
v46May 202463,08619,411171
v47October 202478,72419,433264
v48May 202578,68619,435119
v49September 202578,69119,43313
v50June 202678,73319,44290
Table 4. Human gene counts and coding-set changes across GENCODE v45–v50. We analysed the six comprehensive reference-chromosome GTF files and checked total, coding, and excluded-readthrough counts against each release's published statistics. Entry and exit counts describe identifiers between adjacent rows; they must not be summed to obtain the v45-to-v50 comparison because an identifier can leave and return. Release dates follow GENCODE's release history.[2] Download the release totals and transition counts.

The coding total stayed close to 19,400 even as the broader catalog expanded. The largest step in Table 4 occurs between v46 and v47: all entries increase by 15,638, while coding entries increase by 22. Over the same interval, entries classified specifically as lncRNA rise from 19,258 to 34,914. Changes in an all-gene total can therefore be dominated by RNA annotation rather than protein-coding genes.

For Figure 2, we compiled both counts back to 2014, using the latest published GRCh38 release in each calendar year. The full catalog grows from 60,155 entries in v21 to 78,733 in v50, with a dip in 2016 and a pronounced increase in 2024. The published protein-coding count stays near 20,000, from 19,881 in v21 to 19,442 in v50. The larger total includes non-coding RNA genes, pseudogenes, and other categories.[2][24][21]

Figure 2. All annotated genes and protein-coding genes in GENCODE, 2014–2026. Both lines use the latest published GRCh38 release of each year; 2026 uses v50, released in June and current on September 19. Coding counts follow each release's published definition; the selected releases from 2022 onward exclude readthrough genes. Lines connect annual annotation snapshots, not biological discoveries or losses. We compiled these counts from GENCODE release statistics. Reuse under CC BY 4.0.

We began the timeline in 2014 to keep our comparison within the GRCh38 assembly series. Assembly patch versions change, but every selected statistics page counts only the main reference chromosomes. We used the published summary count for the coding series without retrospectively harmonizing definitions: the selected v42–v50 releases exclude readthrough genes, while v21–v39 do not separately subtract them. The change between the 2021 and 2022 snapshots therefore includes a counting-rule change, not simply a loss of coding genes.[25][26] Annotation rules and evidence also evolve, so these lines describe catalog counts rather than a fixed set of equally validated functional genes. Release months vary, and intermediate releases are omitted. The downloadable chart data records both counts, the coding readthrough policy, release, month, assembly version, source URL, and source-page checksum.

In Figure 3, we grouped identifiers by the annotation changes we observed when they entered or left the coding set between v45 and v50. Of the 60 entering identifiers, 30 were already present under another biotype: 16 as lncRNA and 14 as pseudogenes. One remained classified as protein-coding but lost its readthrough flag, and 29 were absent from the v45 gene records. In the opposite direction, four identifiers changed biotype, three acquired a readthrough flag, and six were absent from v50.

Figure 3. Observed coding-set changes between GENCODE v45 and v50. We analysed reference-chromosome gene records, excluding readthrough genes. Flag changes refer to the readthrough annotation. Bars count identifiers entering or leaving the set, not independently confirmed biological discoveries or losses. Reuse under CC BY 4.0.

Identifier changes require particular care. We found that 24 of the 29 entering identifiers absent from v45 overlap a same-strand gene span already annotated in v45. All six disappearing identifiers overlap another gene span in v50. Overlap alone does not prove that two records represent the same gene, but it shows why “absent identifier” cannot be equated with “previously unknown gene.”

PAXX provides a concrete example: v47 records ENSG00000148362, while v48 records ENSG00000310560 with the same gene name, strand, and genomic span. Identifier matching records an exit and an entry despite this continuity. CAST shows a related complication: its older identifier changes to lncRNA in v47 while another coding identifier with the same name appears at an overlapping location. That change does not establish that CAST ceased to encode a protein.

We counted only GTF gene rows and matched identifiers after removing their numeric version suffix. We preserved any chromosome-Y pseudoautosomal suffix and did not match by gene name.[23] An identifier coding at both endpoints counts as retained even if its transcripts or coordinates changed. We flagged missing identifiers for same-strand span overlap, including introns, without automatically assigning them to splits, mergers, or replacements. We have not adjudicated the underlying cause of every annotation change.

The gene-level comparison, full methods, analysis script, source manifest, and run provenance provide the records, inclusion rules, download URLs, and checksums needed to reproduce the results. Filter the gene-level file to releases 45 and 50 for the endpoint comparison shown in Figure 3.

Does a human cell have 20,000 genes or 40,000?

About 20,000 describes distinct protein-coding gene loci, meaning positions in the genome, rather than all their inherited copies. For most nuclear genes, a person inherits one copy from each parent. Different sequence versions of the same gene are called alleles.[4]

Most human body cells are diploid, with 46 chromosomes arranged in 23 pairs. Sperm and egg cells carry a single chromosome set. Two chromosome sets provide two copies of most genes without doubling the number of distinct gene identities.[5] A gene count is therefore different from a chromosome count: each chromosome contains many genes.

The reference annotation also includes both X and Y and the mitochondrial chromosome, MT. It should not be multiplied by two to obtain an exact gene-copy count for every cell. The guides to genes per chromosome and mitochondrial genes explain these distinct parts of the genome.[3]

Do humans have more genes than other animals?

Humans do not have an exceptionally large protein-coding gene set. Several animal reference annotations contain more coding genes, while others contain fewer. Table 5 retains the Ensembl release snapshots used for this guide rather than treating gene counts as permanent properties of each species.

OrganismCoding genesReference annotation
Human19,442GENCODE v50, GRCh38.p14
Mouse22,081Ensembl 116, GRCm39
Dog20,567Ensembl 116, ROSCfam1.0
Domestic cat19,209Ensembl 116, F.catusFca126mat1.0
Cattle20,848Ensembl 116, ARS-UCD2.0
Zebrafish25,592Ensembl 116, GRCz11
Fruit fly13,986Ensembl 116, BDGP6.46
Arabidopsis thaliana27,655Ensembl Plants 63, TAIR10/Araport11
Baker's yeast6,600Ensembl 116, R64-1-1
Table 5. Protein-coding gene counts in selected species. The human row uses GENCODE. Other rows use the named Ensembl and Ensembl Plants annotations, recorded August 17, 2026.[1][12][13][14][15][16][17][18][19]

Annotation methods and evidence differ among species, so the counts in Table 5 compare the scale of cataloged gene sets rather than provide a standardized measure of biological complexity.

Gene number also differs from the number of gene products. A gene can give rise to several RNA transcripts, and coding transcripts can specify different protein sequences. GENCODE v50 reports 278,455 protein-coding transcripts and 172,117 distinct translations.[1] The translation statistic counts distinct sequences within each gene and sums across genes; it is not a measurement of the proteins present in a particular cell.[3]

The human transcriptome and protein-count guides develop these distinctions. Counting genes describes one level of genome organization; understanding their products also requires knowing which transcripts are expressed, where they are expressed, and how their proteins are processed.

Sources26
  1. Human release statistics (v50)

    GENCODE · September 19, 2026

  2. GENCODE human release history

    GENCODE · September 19, 2026

  3. How GENCODE release statistics are calculated

    GENCODE · September 19, 2026

  4. What is a gene?

    MedlinePlus Genetics, National Library of Medicine · September 19, 2026

  5. Diploid

    National Human Genome Research Institute · September 19, 2026

  6. Gene/transcript biotypes in GENCODE and Ensembl

    GENCODE · September 19, 2026

  7. Homo sapiens Annotation Release GCF_009914755.1-RS_2025_08

    NCBI RefSeq · August 17, 2026

  8. Genenames.org: the HGNC and PGNC resources in 2026

    Nucleic Acids Research · 2026

  9. The state of the human coding gene catalogues

    Database · 2025

  10. 2004: International consortium describes finished human genome sequence

    National Human Genome Research Institute · 2004

  11. 2001: First analysis of the human genome

    National Human Genome Research Institute · 2001

  12. Mouse assembly and gene annotation

    Ensembl · August 17, 2026

  13. Dog assembly and gene annotation

    Ensembl · August 17, 2026

  14. Domestic cat assembly and gene annotation

    Ensembl · August 17, 2026

  15. Cattle assembly and gene annotation

    Ensembl · August 17, 2026

  16. Zebrafish assembly and gene annotation

    Ensembl · August 17, 2026

  17. Fruit fly assembly and gene annotation

    Ensembl · August 17, 2026

  18. Arabidopsis thaliana assembly and gene annotation

    Ensembl Plants · August 17, 2026

  19. Baker’s yeast assembly and gene annotation

    Ensembl · August 17, 2026

  20. Human Release 45 (GRCh38.p14)

    GENCODE · September 19, 2026

  21. Human Release 50 (GRCh38.p14)

    GENCODE · September 19, 2026

  22. GENCODE annotation tags

    GENCODE · September 19, 2026

  23. Format description of GENCODE GTF

    GENCODE · September 19, 2026

  24. Human Release 21 Statistics

    GENCODE · September 19, 2026

  25. Human Release 39 Statistics

    GENCODE · September 19, 2026

  26. Human Release 42 Statistics

    GENCODE · September 19, 2026

Cite this article

Broz, M. (2026, September 19). How many genes do humans have? ProteinIQ. https://proteiniq.io/guides/how-many-genes-do-humans-have

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “How many genes do humans have?” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
January 2, 2026
Updated
September 19, 2026

Related guides

Browse all guides
Ink illustration comparing long and short DNA gene regions with their spans marked.

Genetics · August 10, 2026

What are the largest and smallest human genes?

RBFOX1 is the largest human protein-coding gene by GENCODE v50 genomic span. MLDHR is the shortest under an HGNC-approved protein-coding filter, at 96 base pairs.

Ink illustration of a DNA sequence comparison highlighting a single C-to-T nucleotide difference.

Genetics · July 23, 2026

How many SNPs are in the human genome?

A typical human genome has about 5 million single-nucleotide variants relative to a reference. Roughly 10 million SNP sites are common across the population, while dbSNP Build 157 has 1.172 billion live reference records.

Human and animal illustrations above schematic comparisons of genome alignment, shared genes, and protein sequences.

Genetics · September 25, 2026

How much DNA do humans share with other animals?

Humans and chimpanzees are 98.8% identical across aligned DNA. ProteinIQ's Ensembl analysis of 21 species shows how genome alignment, shared genes, and protein identity diverge with distance.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing