ProteinIQ
Sign inStart for free
ProteinIQ
Genetics

How many CpG sites are in the human genome?

The human genome contains about 28 million CpG sites. A whole-genome methylation atlas counted 28,217,448 positions, while common methylation arrays measure about 1.6% to 3.3% of them.

July 27, 2026·Matic Broz, PhD
Ink illustration of DNA with adjacent cytosine and guanine bases highlighted along one strand.

The human genome contains about 28 million CpG sites. A whole-genome methylation atlas analyzed 28,217,448 CpG positions, equal to roughly one CpG for every 110 base pairs.

These sites are not spread evenly. Most occur at low density, while CpG islands pack many CpGs into short stretches of DNA that often overlap gene-regulatory regions.

How many CpG sites are in the human genome?

One haploid human reference genome contains about 28 million CpG sites. A 2023 atlas used an exact set of 28,217,448 CpG positions.[1]

A CpG site is a cytosine followed immediately by guanine along a DNA strand, with the “p” referring to the phosphate bond between them. Because the sequence is symmetric on the two complementary strands, researchers usually count the genomic position once rather than as two separate CpGs.

Dividing 28,217,448 sites by the roughly 3.1 billion base pairs in the human genome gives an average spacing of about 110 base pairs. The real distribution is highly uneven: CpGs are sparse across much of the genome and concentrated in CpG-rich regions.

The count refers to positions in a reference sequence, not the number of methylated cytosines in a person. A diploid cell carries two copies of most CpG positions, and sequence variants can create or remove individual CpGs.

How many CpG islands are in the human genome?

The human genome has roughly 30,000 CpG islands under a widely used repeat-masked sequence definition, but the count rises above 50,000 when CpG-rich repeats are included.

The initial human genome analysis found 28,890 CpG islands after masking repetitive DNA and 50,267 in the full draft sequence.[2] The large difference came mainly from GC-rich repeat elements, especially Alu sequences.

A CpG island is a region, not one CpG site. The classic sequence rule looks for at least 200 base pairs with GC content above 50% and an observed-to-expected CpG ratio of at least 0.6. The CpG Island Finder applies these criteria to a submitted DNA sequence.

Different minimum lengths, CpG-density thresholds, genome assemblies, and repeat-masking rules produce different totals. “About 30,000” is a useful shorthand only when the counting method is clear.

What percentage of CpG sites do methylation arrays measure?

Common methylation arrays measure about 1.6% to 3.3% of the human genome's CpG sites, depending on the platform.

The 450K array covers roughly 450,000 sites, or 1.6% of 28.2 million. The first MethylationEPIC array covers about 860,000, or 3.0%. Illumina's current MethylationEPIC v2.0 kit targets about 930,000 unique methylation sites, or 3.3%.[1][3]

The whole human genome contains 100% of its roughly 28.2 million CpG sites, while three common methylation arrays cover about 1.6% to 3.3%

The percentages are calculated against the 28,217,448-site set used by the 2023 atlas. Arrays select CpGs in genes, promoters, enhancers, CpG islands, and other regions of interest, so their coverage is not a random 1.6% to 3.3% sample of the genome.

Whole-genome methods can examine far more sites, but the number with usable measurements in a particular experiment depends on sequencing depth, read mapping, sample quality, and analysis filters.

Why are CpG sites rare in the human genome?

CpG sites occur at only about one-fifth of the frequency expected from the human genome's cytosine and guanine content.

Cytosine and guanine each make up about 21% of the genome, so independent pairing would place CpG near 4.4% of dinucleotide positions. The observed frequency is about 0.9%.[2]

The main reason is mutation over evolutionary time. Cytosines in CpG sites are often methylated. When methylated cytosine loses an amino group, it becomes thymine, creating a C-to-T change that DNA repair can miss. Repeated mutations have depleted CpGs from much of the genome.[2]

CpG islands are the main exception. Their CpGs are often unmethylated, which reduces this mutation pressure and helps preserve dense clusters near many gene promoters.

Sources3 references
  1. A DNA methylation atlas of normal human cell types

    Nature · 2023

  2. Initial sequencing and analysis of the human genome

    Nature · 2001

  3. Infinium MethylationEPIC v2.0 Kit

    Illumina · July 27, 2026

About the author

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

LinkedInGoogle ScholarORCID
Published
July 27, 2026
Last updated
July 27, 2026

Related guides

Browse all guides
Nuclear DNA magnified to show the base pairs of a double helix.

Genetics · September 19, 2026

How big is the human genome?

Compare human genome size in base pairs, nucleotides, picograms, and gigabytes, with reference assembly totals and verified download sizes.

Ink illustration of a sample tube, DNA, sequence reads, and a clock representing sequencing turnaround.

Genetics · July 30, 2026

How long does whole-genome sequencing take?

Human whole-genome sequencing can generate genome data in hours to about a day. Sample preparation, analysis, interpretation, and reporting can extend the full turnaround to weeks.

DNA wrapped around histone cores along a nucleosome fiber.

Genetics · July 28, 2026

How many nucleosomes are in a human cell?

A typical diploid human cell contains about 30 million nucleosomes. See how that estimate is calculated and why a nucleosome can be described as containing 146, 166, or about 200 base pairs.

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • Batches
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation
  • For academia
  • For enterprise

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing