ProteinIQ
Sign inStart for free
ProteinIQ
Statistics/Aug '26/9 min read

How many genes do humans have?

Matic Broz

Matic BrozComputational chemist

Humans have about 19,000 protein-coding genes. The current GENCODE human reference annotation gives the exact count as 19,442.

That is the best single answer to the question, but it is not the only valid human gene count. When non-coding RNA genes, pseudogenes, immune-receptor segments, and other annotated features are included, GENCODE lists 78,733 total gene entries.

How many genes do humans have?

GENCODE Release 50 lists 19,442 protein-coding genes on the main human chromosomes. Released in June 2026 for the GRCh38.p14 assembly, it is the current GENCODE human annotation.[1][2]

Human gene categoryGENCODE v50 countWhat the category includes
Protein-coding genes19,442Genes included in GENCODE’s headline protein-coding count
Long non-coding RNA genes35,885Genes annotated as producing long non-coding RNAs
Small non-coding RNA genes7,608miRNA, snoRNA, snRNA, tRNA, and related categories
Pseudogenes14,702Gene-like sequences generally classified separately from functional coding genes
All annotated gene entries78,733Coding genes, non-coding genes, pseudogenes, immune-receptor segments, and other biotypes

These figures come from the GENCODE Release 50 statistics for the main chromosomes.[1]

Figure 1. GENCODE v50 human gene counts by annotation category. The 78,733 total refers to gene entries, not only protein-producing or confirmed functional genes. Source: ProteinIQ analysis of GENCODE Release 50 statistics. Chart may be reused with attribution to ProteinIQ and a link to this page.

The difference between 19,442 and 78,733 is a difference in definition, not an error. The first number counts genes in GENCODE’s main protein-coding category. The larger number counts every annotated gene entry, including genes that do not encode proteins and sequences classified as pseudogenes.

There is a further complication inside the protein-coding category. GENCODE’s detailed biotype table contains 20,107 entries labelled protein_coding, but its headline count excludes 665 readthrough genes:

20,107 protein-coding biotype entries − 665 readthrough genes = 19,442 headline protein-coding genes.

A readthrough gene is annotated where transcription continues across two neighboring loci and produces a combined transcript. Whether such loci should be counted independently is one example of why exact gene totals depend on annotation rules.[1]

The 19,442 figure counts distinct annotated gene loci, not the number of physical gene copies across a person’s body. Most diploid cells carry two copies of most nuclear genes, while eggs and sperm carry one chromosome set. Genes are distributed across the human chromosomes, and the small mitochondrial genome is generally counted separately in discussions of mitochondrial genes.

How many non-coding genes do humans have?

GENCODE Release 50 lists 35,885 long non-coding RNA genes and 7,608 small non-coding RNA genes, giving 43,493 genes across those two headline non-coding RNA categories.[1]

Non-coding RNA genes produce RNA molecules that are not translated into proteins. Depending on the category, these RNAs can help regulate gene expression, process other RNA molecules, assemble cellular machinery, or perform structural and catalytic roles.

Pseudogenes are reported separately. GENCODE lists 14,702 pseudogenes, but it would be misleading to add them to the non-coding RNA total and describe the result simply as “functional non-coding genes.” Some pseudogenes are transcribed or biologically relevant, while others are gene-like evolutionary remnants.

Non-coding genes should also not be confused with all non-coding DNA. Most of the human genome does not encode proteins, but only a subset of that sequence is annotated as belonging to non-coding genes.

Why do GENCODE, RefSeq, and HGNC report different gene counts?

GENCODE, NCBI RefSeq, and HGNC report different human gene counts because they maintain different kinds of scientific catalogs and apply different inclusion rules.

SourceReported protein-coding countWhat is being counted
GENCODE v5019,442Protein-coding genes on the main chromosomes, excluding readthrough genes from the headline total
NCBI RefSeq RS_2025_0819,890Protein-coding genes on the GRCh38 primary assembly
HGNC, September 2025 snapshot19,294Human genes with approved protein-coding gene symbols
2025 three-catalog comparison19,268 consensus genesGenes classified as coding by all three analyzed Ensembl/GENCODE, RefSeq, and UniProtKB catalogs
2025 three-catalog comparison21,871 union genesGenes classified as coding by at least one of the three analyzed catalogs

Download the cross-database human gene count data (CSV)

GENCODE provides genome annotation, RefSeq provides NCBI’s curated and computational reference annotation, and HGNC assigns standardized human gene names and symbols. The cross-catalog study measured where three major coding-gene resources agreed and disagreed.[1][3][4][5]

RefSeq itself illustrates how assembly scope changes the answer. Its RS_2025_08 release reports 19,890 protein-coding genes on the GRCh38 primary assembly, 20,076 across its broader GRCh38 annotation scope, and 20,070 on the telomere-to-telomere T2T-CHM13 assembly.[3]

The 2025 comparison found 2,603 genes whose coding status differed among the three analyzed catalogs. A total of 21,871 genes were considered coding in at least one catalog, but only 19,268 were classified as coding by all three.[5]

That comparison used the database releases available for its analysis, including GENCODE v45, rather than the newer GENCODE v50. Its 19,268 figure is therefore a cross-database consensus snapshot, not a replacement for the current GENCODE count.

Differences can result from:

  • whether readthrough loci are included
  • whether a gene has sufficient evidence that it produces a protein
  • whether pseudogenes or uncertain loci are classified as coding
  • which genome assembly and alternate sequences are included
  • whether the catalog counts genome annotations or approved gene symbols
  • how quickly each database incorporates new evidence

A human gene count is most reproducible when it names the database, release, assembly, and counting rule.

How has the estimated number of human genes changed?

The accepted estimate has fallen from roughly 100,000 genes before the Human Genome Project to about 19,000 protein-coding genes today.

PeriodHuman gene estimateContext
Before the draft human genomeAbout 100,000A widely held expectation before genome-wide sequence analysis
2001 draft genome30,000–35,000Initial analysis of the draft human genome
2004 finished sequence20,000–25,000Revised estimate from the more complete genome sequence
2004 confirmed set19,599Protein-coding genes confirmed in the finished-sequence analysis
GENCODE v50, June 202619,442Current GENCODE headline protein-coding count
Figure 2. Selected historical estimates compared with the current GENCODE v50 annotation count. The first three values are historical estimates from different methods and definitions; the bars do not represent a continuous scientific trend. Source: ProteinIQ analysis of NHGRI historical releases and GENCODE Release 50. Chart may be reused with attribution to ProteinIQ and a link to this page.

The 2001 and 2004 estimates come from Human Genome Project announcements, while the current figure comes from GENCODE Release 50.[6][7][1][2]

In 2004, researchers reported 19,599 confirmed protein-coding genes and another 2,188 predicted coding segments. The estimate continued to change as annotations gained better transcript evidence and researchers reclassified suspected genes as pseudogenes, non-coding loci, readthrough transcripts, or unsupported predictions.[7]

The modern count is not guaranteed to decrease forever. A release can add newly supported genes, remove weak predictions, merge annotations, split loci, or change classifications. That is why a gene count should be treated as a versioned research result rather than a permanent natural constant.

Do humans have more genes than other animals?

Humans do not have an exceptionally large number of protein-coding genes. Several animals and plants have similar or higher coding-gene counts in their current reference annotations.

OrganismCoding genesCurrent reference annotation
Human19,442GENCODE v50
Mouse22,081Ensembl 116, GRCm39
Dog20,567Ensembl 116, ROS_Cfam_1.0
Domestic cat19,209Ensembl 116, F.catus_Fca126_mat1.0
Cattle20,848Ensembl 116, ARS-UCD2.0
Zebrafish25,592Ensembl 116, GRCz11
Fruit fly13,986Ensembl 116, BDGP6.46
Arabidopsis thaliana27,655Ensembl Plants 63, TAIR10/Araport11
Baker’s yeast6,600Ensembl 116, R64-1-1

The human count comes from GENCODE Release 50. The animal, fruit-fly, and yeast figures come from Ensembl Release 116, while the Arabidopsis count comes from Ensembl Plants Release 63.[1][8][9][10][11][12][13][14][15]

These are reference-annotation counts, not perfectly standardized measurements. Different annotation projects can use different assemblies, evidence thresholds, and gene classifications. The table is therefore most useful for comparing broad scale rather than claiming that one species has precisely a certain number more “real genes” than another.

Gene number is not a simple measure of biological complexity. Humans derive complexity from when and where genes are expressed, alternative transcripts, protein processing, regulatory networks, cell specialization, and interactions among genes and proteins.

One gene can produce multiple transcripts, and those transcripts can lead to different translated products. GENCODE v50 contains 644,292 total transcripts, including 278,455 protein-coding transcripts and 172,117 distinct translations, despite having only 19,442 headline protein-coding genes.[1]

The distinctions between genes, transcripts, translations, and physical protein molecules are explored further in the guides to the human transcriptome and the number of proteins. Gene count is also different from DNA similarity between humans and other animals.

Which human gene count should researchers report?

For a current general reference, report 19,442 protein-coding genes and identify the source as GENCODE Release 50 on GRCh38.p14, released in June 2026.[1][2]

Use 78,733 total annotated gene entries only when the intended count includes non-coding RNA genes, pseudogenes, immune-receptor segments, and other annotated gene categories. Do not present 78,733 as the number of human protein-coding or necessarily functional genes.

Use 19,268 consensus protein-coding genes when discussing the 2025 analysis of agreement among Ensembl/GENCODE, RefSeq, and UniProtKB. That number describes the intersection of the catalog versions analyzed in that study, while 21,871 describes their union.[5]

A concise, definition-safe summary is:

GENCODE Release 50, released in June 2026, annotates 19,442 protein-coding genes and 78,733 total gene entries on the main human chromosomes.

Because genome annotations continue to change, any publication, dataset, or article reporting a precise human gene count should include the database name and release. A bare statement such as “humans have exactly 20,000 genes” is easy to understand but cannot be independently reproduced without knowing what was counted.

Sources▼
  1. Human release statistics (v50) GENCODE · August 17, 2026. https://www.gencodegenes.org/human/stats.html
  2. GENCODE human release history GENCODE · August 17, 2026. https://www.gencodegenes.org/human/releases.html
  3. Homo sapiens Annotation Release GCF_009914755.1-RS_2025_08 NCBI RefSeq · August 17, 2026. https://www.ncbi.nlm.nih.gov/refseq/annotation_euk/Homo_sapiens/GCF_009914755.1-RS_2025_08/
  4. Genenames.org: the HGNC and PGNC resources in 2026 Nucleic Acids Research · 2026. https://pmc.ncbi.nlm.nih.gov/articles/PMC12807706/
  5. The state of the human coding gene catalogues Database · 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12462614/
  6. 2001: First analysis of the human genome National Human Genome Research Institute · 2001. https://www.genome.gov/10002192/2001-release-first-analysis-of-human-genome
  7. 2004: International consortium describes finished human genome sequence National Human Genome Research Institute · 2004. https://www.genome.gov/12513430/2004-release-ihgsc-describes-finished-human-sequence
  8. Mouse assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Mus_musculus/Info/Annotation
  9. Dog assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Canis_lupus_familiaris/Info/Annotation
  10. Domestic cat assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Felis_catus/Info/Annotation
  11. Cattle assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Bos_taurus/Info/Annotation
  12. Zebrafish assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Danio_rerio/Info/Annotation
  13. Fruit fly assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Drosophila_melanogaster/Info/Annotation
  14. Arabidopsis thaliana assembly and gene annotation Ensembl Plants · August 17, 2026. https://plants.ensembl.org/Arabidopsis_thaliana/Info/Annotation/
  15. Baker’s yeast assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Saccharomyces_cerevisiae/Info/Annotation
Published
January 2, 2026
Last updated
August 17, 2026

Table of contents

Cite this article

Broz, M. (2026, August 17). How many genes do humans have? ProteinIQ. https://proteiniq.io/guides/how-many-genes-do-humans-have

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

Related guides

Statistics

AUG '26

What are the largest and smallest human genes?

RBFOX1 is the largest human protein-coding gene by GENCODE v50 genomic span. MLDHR is the shortest under an HGNC-approved protein-coding filter, at 96 base pairs.

Matic Broz Computational chemist

Statistics

JUL '26

How many SNPs are in the human genome?

A typical human genome has about 5 million single-nucleotide variants relative to a reference. Roughly 10 million SNP sites are common across the population, while dbSNP Build 157 has 1.172 billion live reference records.

Matic Broz Computational chemist

Statistics

JUL '26

How many genes are in mitochondrial DNA?

Human mitochondrial DNA contains 37 genes in a circular genome of 16,569 base pairs. Most human protein-coding genes are instead found in nuclear DNA.

Matic Broz Computational chemist

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing