Humans have about 19,000 protein-coding genes. The current GENCODE human reference annotation gives the exact count as 19,442.
That is the best single answer to the question, but it is not the only valid human gene count. When non-coding RNA genes, pseudogenes, immune-receptor segments, and other annotated features are included, GENCODE lists 78,733 total gene entries.
How many genes do humans have?
GENCODE Release 50 lists 19,442 protein-coding genes on the main human chromosomes. Released in June 2026 for the GRCh38.p14 assembly, it is the current GENCODE human annotation.[1][2]
| Human gene category | GENCODE v50 count | What the category includes |
|---|---|---|
| Protein-coding genes | 19,442 | Genes included in GENCODE’s headline protein-coding count |
| Long non-coding RNA genes | 35,885 | Genes annotated as producing long non-coding RNAs |
| Small non-coding RNA genes | 7,608 | miRNA, snoRNA, snRNA, tRNA, and related categories |
| Pseudogenes | 14,702 | Gene-like sequences generally classified separately from functional coding genes |
| All annotated gene entries | 78,733 | Coding genes, non-coding genes, pseudogenes, immune-receptor segments, and other biotypes |
These figures come from the GENCODE Release 50 statistics for the main chromosomes.[1]
The difference between 19,442 and 78,733 is a difference in definition, not an error. The first number counts genes in GENCODE’s main protein-coding category. The larger number counts every annotated gene entry, including genes that do not encode proteins and sequences classified as pseudogenes.
There is a further complication inside the protein-coding category. GENCODE’s detailed biotype table contains 20,107 entries labelled protein_coding, but its headline count excludes 665 readthrough genes:
20,107 protein-coding biotype entries − 665 readthrough genes = 19,442 headline protein-coding genes.
A readthrough gene is annotated where transcription continues across two neighboring loci and produces a combined transcript. Whether such loci should be counted independently is one example of why exact gene totals depend on annotation rules.[1]
The 19,442 figure counts distinct annotated gene loci, not the number of physical gene copies across a person’s body. Most diploid cells carry two copies of most nuclear genes, while eggs and sperm carry one chromosome set. Genes are distributed across the human chromosomes, and the small mitochondrial genome is generally counted separately in discussions of mitochondrial genes.
How many non-coding genes do humans have?
GENCODE Release 50 lists 35,885 long non-coding RNA genes and 7,608 small non-coding RNA genes, giving 43,493 genes across those two headline non-coding RNA categories.[1]
Non-coding RNA genes produce RNA molecules that are not translated into proteins. Depending on the category, these RNAs can help regulate gene expression, process other RNA molecules, assemble cellular machinery, or perform structural and catalytic roles.
Pseudogenes are reported separately. GENCODE lists 14,702 pseudogenes, but it would be misleading to add them to the non-coding RNA total and describe the result simply as “functional non-coding genes.” Some pseudogenes are transcribed or biologically relevant, while others are gene-like evolutionary remnants.
Non-coding genes should also not be confused with all non-coding DNA. Most of the human genome does not encode proteins, but only a subset of that sequence is annotated as belonging to non-coding genes.
Why do GENCODE, RefSeq, and HGNC report different gene counts?
GENCODE, NCBI RefSeq, and HGNC report different human gene counts because they maintain different kinds of scientific catalogs and apply different inclusion rules.
| Source | Reported protein-coding count | What is being counted |
|---|---|---|
| GENCODE v50 | 19,442 | Protein-coding genes on the main chromosomes, excluding readthrough genes from the headline total |
| NCBI RefSeq RS_2025_08 | 19,890 | Protein-coding genes on the GRCh38 primary assembly |
| HGNC, September 2025 snapshot | 19,294 | Human genes with approved protein-coding gene symbols |
| 2025 three-catalog comparison | 19,268 consensus genes | Genes classified as coding by all three analyzed Ensembl/GENCODE, RefSeq, and UniProtKB catalogs |
| 2025 three-catalog comparison | 21,871 union genes | Genes classified as coding by at least one of the three analyzed catalogs |
Download the cross-database human gene count data (CSV)
GENCODE provides genome annotation, RefSeq provides NCBI’s curated and computational reference annotation, and HGNC assigns standardized human gene names and symbols. The cross-catalog study measured where three major coding-gene resources agreed and disagreed.[1][3][4][5]
RefSeq itself illustrates how assembly scope changes the answer. Its RS_2025_08 release reports 19,890 protein-coding genes on the GRCh38 primary assembly, 20,076 across its broader GRCh38 annotation scope, and 20,070 on the telomere-to-telomere T2T-CHM13 assembly.[3]
The 2025 comparison found 2,603 genes whose coding status differed among the three analyzed catalogs. A total of 21,871 genes were considered coding in at least one catalog, but only 19,268 were classified as coding by all three.[5]
That comparison used the database releases available for its analysis, including GENCODE v45, rather than the newer GENCODE v50. Its 19,268 figure is therefore a cross-database consensus snapshot, not a replacement for the current GENCODE count.
Differences can result from:
- whether readthrough loci are included
- whether a gene has sufficient evidence that it produces a protein
- whether pseudogenes or uncertain loci are classified as coding
- which genome assembly and alternate sequences are included
- whether the catalog counts genome annotations or approved gene symbols
- how quickly each database incorporates new evidence
A human gene count is most reproducible when it names the database, release, assembly, and counting rule.
How has the estimated number of human genes changed?
The accepted estimate has fallen from roughly 100,000 genes before the Human Genome Project to about 19,000 protein-coding genes today.
| Period | Human gene estimate | Context |
|---|---|---|
| Before the draft human genome | About 100,000 | A widely held expectation before genome-wide sequence analysis |
| 2001 draft genome | 30,000–35,000 | Initial analysis of the draft human genome |
| 2004 finished sequence | 20,000–25,000 | Revised estimate from the more complete genome sequence |
| 2004 confirmed set | 19,599 | Protein-coding genes confirmed in the finished-sequence analysis |
| GENCODE v50, June 2026 | 19,442 | Current GENCODE headline protein-coding count |
The 2001 and 2004 estimates come from Human Genome Project announcements, while the current figure comes from GENCODE Release 50.[6][7][1][2]
In 2004, researchers reported 19,599 confirmed protein-coding genes and another 2,188 predicted coding segments. The estimate continued to change as annotations gained better transcript evidence and researchers reclassified suspected genes as pseudogenes, non-coding loci, readthrough transcripts, or unsupported predictions.[7]
The modern count is not guaranteed to decrease forever. A release can add newly supported genes, remove weak predictions, merge annotations, split loci, or change classifications. That is why a gene count should be treated as a versioned research result rather than a permanent natural constant.
Do humans have more genes than other animals?
Humans do not have an exceptionally large number of protein-coding genes. Several animals and plants have similar or higher coding-gene counts in their current reference annotations.
| Organism | Coding genes | Current reference annotation |
|---|---|---|
| Human | 19,442 | GENCODE v50 |
| Mouse | 22,081 | Ensembl 116, GRCm39 |
| Dog | 20,567 | Ensembl 116, ROS_Cfam_1.0 |
| Domestic cat | 19,209 | Ensembl 116, F.catus_Fca126_mat1.0 |
| Cattle | 20,848 | Ensembl 116, ARS-UCD2.0 |
| Zebrafish | 25,592 | Ensembl 116, GRCz11 |
| Fruit fly | 13,986 | Ensembl 116, BDGP6.46 |
| Arabidopsis thaliana | 27,655 | Ensembl Plants 63, TAIR10/Araport11 |
| Baker’s yeast | 6,600 | Ensembl 116, R64-1-1 |
The human count comes from GENCODE Release 50. The animal, fruit-fly, and yeast figures come from Ensembl Release 116, while the Arabidopsis count comes from Ensembl Plants Release 63.[1][8][9][10][11][12][13][14][15]
These are reference-annotation counts, not perfectly standardized measurements. Different annotation projects can use different assemblies, evidence thresholds, and gene classifications. The table is therefore most useful for comparing broad scale rather than claiming that one species has precisely a certain number more “real genes” than another.
Gene number is not a simple measure of biological complexity. Humans derive complexity from when and where genes are expressed, alternative transcripts, protein processing, regulatory networks, cell specialization, and interactions among genes and proteins.
One gene can produce multiple transcripts, and those transcripts can lead to different translated products. GENCODE v50 contains 644,292 total transcripts, including 278,455 protein-coding transcripts and 172,117 distinct translations, despite having only 19,442 headline protein-coding genes.[1]
The distinctions between genes, transcripts, translations, and physical protein molecules are explored further in the guides to the human transcriptome and the number of proteins. Gene count is also different from DNA similarity between humans and other animals.
Which human gene count should researchers report?
For a current general reference, report 19,442 protein-coding genes and identify the source as GENCODE Release 50 on GRCh38.p14, released in June 2026.[1][2]
Use 78,733 total annotated gene entries only when the intended count includes non-coding RNA genes, pseudogenes, immune-receptor segments, and other annotated gene categories. Do not present 78,733 as the number of human protein-coding or necessarily functional genes.
Use 19,268 consensus protein-coding genes when discussing the 2025 analysis of agreement among Ensembl/GENCODE, RefSeq, and UniProtKB. That number describes the intersection of the catalog versions analyzed in that study, while 21,871 describes their union.[5]
A concise, definition-safe summary is:
GENCODE Release 50, released in June 2026, annotates 19,442 protein-coding genes and 78,733 total gene entries on the main human chromosomes.
Because genome annotations continue to change, any publication, dataset, or article reporting a precise human gene count should include the database name and release. A bare statement such as “humans have exactly 20,000 genes” is easy to understand but cannot be independently reproduced without knowing what was counted.
Sources▼
- Human release statistics (v50) GENCODE · August 17, 2026. https://www.gencodegenes.org/human/stats.html
- GENCODE human release history GENCODE · August 17, 2026. https://www.gencodegenes.org/human/releases.html
- Homo sapiens Annotation Release GCF_009914755.1-RS_2025_08 NCBI RefSeq · August 17, 2026. https://www.ncbi.nlm.nih.gov/refseq/annotation_euk/Homo_sapiens/GCF_009914755.1-RS_2025_08/
- Genenames.org: the HGNC and PGNC resources in 2026 Nucleic Acids Research · 2026. https://pmc.ncbi.nlm.nih.gov/articles/PMC12807706/
- The state of the human coding gene catalogues Database · 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12462614/
- 2001: First analysis of the human genome National Human Genome Research Institute · 2001. https://www.genome.gov/10002192/2001-release-first-analysis-of-human-genome
- 2004: International consortium describes finished human genome sequence National Human Genome Research Institute · 2004. https://www.genome.gov/12513430/2004-release-ihgsc-describes-finished-human-sequence
- Mouse assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Mus_musculus/Info/Annotation
- Dog assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Canis_lupus_familiaris/Info/Annotation
- Domestic cat assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Felis_catus/Info/Annotation
- Cattle assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Bos_taurus/Info/Annotation
- Zebrafish assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Danio_rerio/Info/Annotation
- Fruit fly assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Drosophila_melanogaster/Info/Annotation
- Arabidopsis thaliana assembly and gene annotation Ensembl Plants · August 17, 2026. https://plants.ensembl.org/Arabidopsis_thaliana/Info/Annotation/
- Baker’s yeast assembly and gene annotation Ensembl · August 17, 2026. https://www.ensembl.org/Saccharomyces_cerevisiae/Info/Annotation

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.