ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Genetics

How big is the human genome?

September 19, 2026·Matic Broz, PhD
Nuclear DNA magnified to show the base pairs of a double helix.

The human genome contains about 3.1 billion base pairs per chromosome set, or roughly 6.3 billion in a typical diploid nucleus before DNA replication. The first figure describes a haploid genome; the second includes the chromosome sets inherited from both parents.

These are useful approximations, rather than exact counts shared by every person. More precise answers depend on the chromosomes included, the reference sequence used, and whether “size” means DNA sequence length, molecular mass, or computer storage.

How many base pairs and nucleotides are in the human genome?

A haploid human genome has about 3.1 billion base pairs, equivalent to about 6.2 billion nucleotides across both strands of its DNA. Each base pair joins two nucleotides, one on each strand: A pairs with T, and C pairs with G.[1]

The two DNA strands are complementary parts of the same chromosome, not the two chromosome sets in a diploid cell. This distinction separates two different reasons for doubling a count: counting both strands doubles the nucleotide total, while counting both parental chromosome sets gives the diploid DNA content.

Figure 1. Base pairs and nucleotides count different units. Nucleotides are counted across both DNA strands. Diploid values describe nuclear DNA before replication; all values are rounded. Reuse under CC BY 4.0.
DNA contentBase pairsNucleotides across both strands
One haploid chromosome setAbout 3.1 billionAbout 6.2 billion
Diploid nucleus before replicationAbout 6.3 billionAbout 12.6 billion

The base-pair figures summarize the conventional haploid estimate and published diploid estimates of 6.27–6.37 billion bp. We calculated the nucleotide totals by doubling the rounded base-pair counts. The haploid and diploid figures are independently rounded, so doubling 3.1 does not reproduce the displayed 6.3.[1][2]

Sequence files normally record one strand at each position, since the complementary strand can be inferred. Consequently, a reference described as containing 3.1 billion bases or nucleotide positions does not imply that the double-stranded molecules contain only 3.1 billion physical nucleotides.

How much DNA is in a human cell?

A typical human diploid nucleus contains approximately 6.4–6.5 picograms of DNA before replication. One picogram is one trillionth of a gram. Piovesan and colleagues calculated the following values from GRCh38.p10, including estimates for uncertain sequence.[2]

Reference chromosome complementNuclear sequence lengthCalculated nuclear DNA mass
46,XY6.27 billion bp6.41 pg
46,XX6.37 billion bp6.51 pg

These are the study's reference-based calculations, not measurements of every individual. The larger X chromosome accounts for the higher 46,XX total. Both rows exclude mitochondrial DNA.[2]

DNA content also changes during the cell cycle. Replication doubles the amount of DNA before division, even though the cell remains diploid: each chromosome now consists of two sister chromatids. Ploidy describes chromosome sets, whereas DNA content also depends on whether those chromosomes have been copied.[3]

Sperm contain one chromosome set. An ovulated human egg also has a haploid chromosome complement, but its chromosomes still have two sister chromatids until meiosis II is completed following fertilization. It therefore needs a stage-specific DNA count.[3] Mature red blood cells lack nuclear DNA, and mitochondrial copy number varies among cell types.[2]

Sequence length and mass are separate from the physical length of DNA. The familiar estimate of about two meters per diploid nucleus describes the DNA extended along its molecular axis, rather than the space it occupies inside the cell.

Why do reference genomes have different sizes?

A reference assembly is a defined collection of sequences used as a coordinate system. Its total need not equal the DNA in one person's haploid chromosome set. GRCh38, for example, includes both X and Y; an ordinary haploid set contains one of these sex chromosomes.[4]

Reference and counting ruleReported sequence length
GRCh38.p14, GRC all-scaffold total including gaps3,099,734,149 bp
GRCh38.p14, placed scaffolds including gaps3,088,269,832 bp
GRCh38.p14, all-scaffold ungapped total2,948,611,470 bp
T2T-CHM13, 2022 paper, chromosomes 1–22 and X3.055 billion bp

The first three rows reproduce the Genome Reference Consortium's reported totals. A scaffold is an assembled sequence that can include unresolved gaps; the ungapped count excludes those gaps. The final row is the total reported by Nurk and colleagues for T2T-CHM13.[4][5]

T2T-CHM13 resolved repetitive regions missing from earlier references, adding nearly 200 million base pairs of sequence. Its smaller headline total does not contradict that gain: GRCh38's gapped length includes estimated gaps, and the two assemblies differ in chromosome content and sequence composition.[5] The history of human genome sequencing therefore concerns completeness as well as total length.

The 2022 CHM13 assembly lacked Y. In 2023, the T2T consortium published a complete 62,460,029-bp Y chromosome from a different individual, HG002. This is a separate chromosome assembly, not evidence that the original CHM13 genome contained Y.[6]

Download bundles can also include alternative versions of regions and patches. Adding every sequence in such a bundle counts some loci more than once. For a reproducible comparison, name the assembly release, sequence set, and treatment of gaps.[7] Genome length should also be distinguished from gene count, which depends on annotation.

How large is the human genome as a computer file?

About 3.1 billion DNA letters require about 3.1 GB of storage at one byte per letter, before file-format overhead. This is a representation of a reference sequence, not the size of all sequencing data collected from a person.

Genomic and storage units use different conventions. A kilobase (kb) means 1,000 bases, a megabase (Mb) means one million, and a gigabase (Gb) means one billion. GB denotes a billion bytes; GiB denotes 1,073,741,824 bytes. In computing, lowercase b can instead mean bits, so the context matters.

Unit or representationEquivalent for 3,099,734,149 sequence positions
Kilobases3,099,734.149 kb
Megabases3,099.734149 Mb
GigabasesAbout 3.10 Gb
One-byte-per-letter textAbout 3.10 GB, or 2.89 GiB
Two-bit A/C/G/T encoding, before metadataAbout 0.775 GB

We converted the GRC all-scaffold total into these units. The storage rows are calculations, not measured downloads.[4] FASTA adds sequence headers and usually line breaks to the letters.[8] The two-bit row assumes four possible bases; an actual .2bit file additionally records sequence names, unknown-base regions, masking, and indexing information.[9]

Actual reference downloads depend on both encoding and sequence content. The following are the original hg38 files in UCSC's bigZips/ directory, which its documentation identifies with the initial 2013 release. They are not the patched files in the latest/ directory.

UCSC reference fileReported file lengthDecimal size
hg38.2bit835,393,456 bytesAbout 835 MB, or 0.835 GB
hg38.fa.gz983,659,424 bytesAbout 984 MB, or 0.984 GB

File lengths are the server's Content-Length values checked on September 19, 2026; we converted bytes to decimal MB and GB. The directory displays these as 797M and 938M, respectively, using rounded binary-scale values. These downloads include sequences beyond the main chromosomes, so they are not an encoding-only comparison with the GRC total above.[7]

Whole-genome sequencing produces much more data because it reads positions repeatedly. For example, Illumina's NovaSeq X throughput estimates assume more than 120 Gb of sequence per sample to achieve 30× genome coverage. That is an instrument-planning assumption expressed in bases, not a 120 GB file-size specification.[10] FASTQ files additionally store read identifiers and quality scores, and compression changes their storage requirements.[11] A useful sequencing storage estimate must therefore specify coverage, file format, and compression, rather than genome length alone.

Sources11
  1. Base Pair

    National Human Genome Research Institute · September 19, 2026

  2. On the length, weight and GC content of the human genome

    BMC Research Notes · 2019

  3. Meiosis and Fertilization

    The Cell: A Molecular Approach, NCBI Bookshelf · 2000

  4. Human Genome Assembly GRCh38.p14

    Genome Reference Consortium · September 19, 2026

  5. The complete sequence of a human genome

    Science · 2022

  6. The complete sequence of a human Y chromosome

    Nature · 2023

  7. Human hg38 downloads and assembly documentation

    UCSC Genome Browser · September 19, 2026

  8. FASTA Format for Nucleotide Sequences

    NCBI GenBank · September 19, 2026

  9. Genome Browser file formats: .2bit format

    UCSC Genome Browser · September 19, 2026

  10. NovaSeq X Specifications

    Illumina · September 19, 2026

  11. FastQ Files

    Illumina Connected Software · September 19, 2026

Cite this article

Broz, M. (2026, September 19). How big is the human genome? ProteinIQ. https://proteiniq.io/guides/human-genome-size

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “How big is the human genome?” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
July 1, 2026
Updated
September 19, 2026

Related guides

Browse all guides
DNA and sequencing-read motifs beside coins representing genome sequencing cost.

Genetics · September 24, 2026

How much does it cost to sequence a genome?

A human genome costs about $220 to $450 at university sequencing labs and $399 to $595 as a consumer test. ProteinIQ's survey of US core facilities puts the median lab price at $419 before analysis.

A magnifying glass examines DNA alongside overlapping sequence fragments.

Genetics · August 10, 2026

How accurate is DNA sequencing?

DNA sequencing accuracy ranges from about 99% per raw base to above 99.9% for high-quality or consensus reads. The exact figure depends on the platform, chemistry, software, sample, and metric.

Ink illustration of a sample tube, DNA, sequence reads, and a clock representing sequencing turnaround.

Genetics · July 30, 2026

How long does whole-genome sequencing take?

Human whole-genome sequencing can generate genome data in hours to about a day. Sample preparation, analysis, interpretation, and reporting can extend the full turnaround to weeks.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing