
The human genome contains about 3.1 billion base pairs per chromosome set, or roughly 6.3 billion in a typical diploid nucleus before DNA replication. The first figure describes a haploid genome; the second includes the chromosome sets inherited from both parents.
These are useful approximations, rather than exact counts shared by every person. More precise answers depend on the chromosomes included, the reference sequence used, and whether “size” means DNA sequence length, molecular mass, or computer storage.
How many base pairs and nucleotides are in the human genome?
A haploid human genome has about 3.1 billion base pairs, equivalent to about 6.2 billion nucleotides across both strands of its DNA. Each base pair joins two nucleotides, one on each strand: A pairs with T, and C pairs with G.[1]
The two DNA strands are complementary parts of the same chromosome, not the two chromosome sets in a diploid cell. This distinction separates two different reasons for doubling a count: counting both strands doubles the nucleotide total, while counting both parental chromosome sets gives the diploid DNA content.
| DNA content | Base pairs | Nucleotides across both strands |
|---|---|---|
| One haploid chromosome set | About 3.1 billion | About 6.2 billion |
| Diploid nucleus before replication | About 6.3 billion | About 12.6 billion |
The base-pair figures summarize the conventional haploid estimate and published diploid estimates of 6.27–6.37 billion bp. We calculated the nucleotide totals by doubling the rounded base-pair counts. The haploid and diploid figures are independently rounded, so doubling 3.1 does not reproduce the displayed 6.3.[1][2]
Sequence files normally record one strand at each position, since the complementary strand can be inferred. Consequently, a reference described as containing 3.1 billion bases or nucleotide positions does not imply that the double-stranded molecules contain only 3.1 billion physical nucleotides.
How much DNA is in a human cell?
A typical human diploid nucleus contains approximately 6.4–6.5 picograms of DNA before replication. One picogram is one trillionth of a gram. Piovesan and colleagues calculated the following values from GRCh38.p10, including estimates for uncertain sequence.[2]
| Reference chromosome complement | Nuclear sequence length | Calculated nuclear DNA mass |
|---|---|---|
| 46,XY | 6.27 billion bp | 6.41 pg |
| 46,XX | 6.37 billion bp | 6.51 pg |
These are the study's reference-based calculations, not measurements of every individual. The larger X chromosome accounts for the higher 46,XX total. Both rows exclude mitochondrial DNA.[2]
DNA content also changes during the cell cycle. Replication doubles the amount of DNA before division, even though the cell remains diploid: each chromosome now consists of two sister chromatids. Ploidy describes chromosome sets, whereas DNA content also depends on whether those chromosomes have been copied.[3]
Sperm contain one chromosome set. An ovulated human egg also has a haploid chromosome complement, but its chromosomes still have two sister chromatids until meiosis II is completed following fertilization. It therefore needs a stage-specific DNA count.[3] Mature red blood cells lack nuclear DNA, and mitochondrial copy number varies among cell types.[2]
Sequence length and mass are separate from the physical length of DNA. The familiar estimate of about two meters per diploid nucleus describes the DNA extended along its molecular axis, rather than the space it occupies inside the cell.
Why do reference genomes have different sizes?
A reference assembly is a defined collection of sequences used as a coordinate system. Its total need not equal the DNA in one person's haploid chromosome set. GRCh38, for example, includes both X and Y; an ordinary haploid set contains one of these sex chromosomes.[4]
| Reference and counting rule | Reported sequence length |
|---|---|
| GRCh38.p14, GRC all-scaffold total including gaps | 3,099,734,149 bp |
| GRCh38.p14, placed scaffolds including gaps | 3,088,269,832 bp |
| GRCh38.p14, all-scaffold ungapped total | 2,948,611,470 bp |
| T2T-CHM13, 2022 paper, chromosomes 1–22 and X | 3.055 billion bp |
The first three rows reproduce the Genome Reference Consortium's reported totals. A scaffold is an assembled sequence that can include unresolved gaps; the ungapped count excludes those gaps. The final row is the total reported by Nurk and colleagues for T2T-CHM13.[4][5]
T2T-CHM13 resolved repetitive regions missing from earlier references, adding nearly 200 million base pairs of sequence. Its smaller headline total does not contradict that gain: GRCh38's gapped length includes estimated gaps, and the two assemblies differ in chromosome content and sequence composition.[5] The history of human genome sequencing therefore concerns completeness as well as total length.
The 2022 CHM13 assembly lacked Y. In 2023, the T2T consortium published a complete 62,460,029-bp Y chromosome from a different individual, HG002. This is a separate chromosome assembly, not evidence that the original CHM13 genome contained Y.[6]
Download bundles can also include alternative versions of regions and patches. Adding every sequence in such a bundle counts some loci more than once. For a reproducible comparison, name the assembly release, sequence set, and treatment of gaps.[7] Genome length should also be distinguished from gene count, which depends on annotation.
How large is the human genome as a computer file?
About 3.1 billion DNA letters require about 3.1 GB of storage at one byte per letter, before file-format overhead. This is a representation of a reference sequence, not the size of all sequencing data collected from a person.
Genomic and storage units use different conventions. A kilobase (kb) means 1,000 bases, a megabase (Mb) means one million, and a gigabase (Gb) means one billion. GB denotes a billion bytes; GiB denotes 1,073,741,824 bytes. In computing, lowercase b can instead mean bits, so the context matters.
| Unit or representation | Equivalent for 3,099,734,149 sequence positions |
|---|---|
| Kilobases | 3,099,734.149 kb |
| Megabases | 3,099.734149 Mb |
| Gigabases | About 3.10 Gb |
| One-byte-per-letter text | About 3.10 GB, or 2.89 GiB |
| Two-bit A/C/G/T encoding, before metadata | About 0.775 GB |
We converted the GRC all-scaffold total into these units. The storage rows are calculations, not measured downloads.[4] FASTA adds sequence headers and usually line breaks to the letters.[8] The two-bit row assumes four possible bases; an actual .2bit file additionally records sequence names, unknown-base regions, masking, and indexing information.[9]
Actual reference downloads depend on both encoding and sequence content. The following are the original hg38 files in UCSC's bigZips/ directory, which its documentation identifies with the initial 2013 release. They are not the patched files in the latest/ directory.
| UCSC reference file | Reported file length | Decimal size |
|---|---|---|
hg38.2bit | 835,393,456 bytes | About 835 MB, or 0.835 GB |
hg38.fa.gz | 983,659,424 bytes | About 984 MB, or 0.984 GB |
File lengths are the server's Content-Length values checked on September 19, 2026; we converted bytes to decimal MB and GB. The directory displays these as 797M and 938M, respectively, using rounded binary-scale values. These downloads include sequences beyond the main chromosomes, so they are not an encoding-only comparison with the GRC total above.[7]
Whole-genome sequencing produces much more data because it reads positions repeatedly. For example, Illumina's NovaSeq X throughput estimates assume more than 120 Gb of sequence per sample to achieve 30× genome coverage. That is an instrument-planning assumption expressed in bases, not a 120 GB file-size specification.[10] FASTQ files additionally store read identifiers and quality scores, and compression changes their storage requirements.[11] A useful sequencing storage estimate must therefore specify coverage, file format, and compression, rather than genome length alone.


