ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Genetics

What percentage of the human genome is viral DNA?

July 23, 2026·Matic Broz, PhD
Ink illustration of a highlighted DNA segment representing inherited sequences of ancient viral origin.

About 8% of the human genome is viral DNA in the usual sense of the question: ancient retroviral sequences that entered the germline and became inherited human DNA.

That is roughly 250 million bases in one haploid human reference genome. It is more than six times the DNA that directly codes for proteins, but it is not evidence that people carry active viruses in one-twelfth of their genome.

What percentage of the human genome is viral DNA?

The best single answer is about 8%. This is the fraction usually attributed to human endogenous retroviruses, or HERVs: DNA sequences left by retroviruses that infected the germline of our primate ancestors.

A 2016 HERV classification paper states that HERVs constitute 8% of the human genome and notes that this includes many single long terminal repeats and older defective LTR retroelement fragments. A 2004 PNAS study used the older 5-8% range and counted about 98,000 HERV elements and fragments.[1][2]

The current GRCh38.p14 reference assembly has 3,099,734,149 base pairs across all scaffolds. Multiplying that by 8% gives 247,978,732 base pairs, so the scale is about 250 million bases of viral-origin sequence per haploid genome.[3]

For comparison, ENCODE reported that protein-coding exons cover 1.22% of the genome. On that basis, recognizable viral-origin sequence occupies about six to seven times as much DNA as the sequence that directly encodes proteins.[4]

Why do some sources give a number closer to half?

Numbers close to half of the genome usually refer to repetitive or mobile DNA, not strictly viral DNA. RepeatMasker reports that 48.49% of hg38 is masked as interspersed repeats, and 52.58% is masked when simple repeats, tandem repeats, satellite DNA, and low-complexity regions are included.[5]

That broader repeat category includes several families with different histories: LINEs, SINEs, LTR retroelements, DNA transposons, satellites, and simple repeats. The LTR retroelement group includes endogenous retroviruses, but LINEs and SINEs should not be counted as ancient viruses in the same plain-language sense.

This is why "8% viral DNA" and "about half repetitive DNA" can both be true. They answer different questions. The first is about recognizable retroviral-origin sequence. The second is about the wider repeat landscape of the human genome, much of which is discussed as transposable or repetitive DNA rather than viral DNA.

Which sequences in the human genome are derived from viruses?

Most viral-derived DNA in the human genome is made of endogenous retrovirus sequences: long terminal repeats, damaged proviral fragments, and occasional more complete proviruses with recognizable gag, pol, or env parts.

A full retrovirus-like provirus has long terminal repeats at both ends and internal genes related to retroviral replication. Over time, most HERV copies lost protein-coding capacity through point mutations, frameshifts, deletions, and recombination that left behind solo LTRs.[1]

Non-retroviral viral fossils also exist, but they are much rarer and are not what people usually mean by the 8% statistic. Katzourakis and Gifford identified endogenous viral elements from ten non-retroviral families in animal genomes, including RNA and DNA virus groups; retroviruses dominate because integration into host DNA is part of their normal replication cycle.[6]

The practical split is simple:

Sequence classScaleDescription
Human endogenous retroviruses and related LTR elementsAbout 8% of the genomeThe standard answer for viral-origin human DNA
All interspersed repeats in hg3848.49% of the genomeRepetitive/mobile DNA, not all viral DNA
Protein-coding exons1.22% of the genomeDNA that directly encodes proteins
Viral-origin DNA compared with other human genome fractions, including coding exons, HERV and LTR elements, interspersed repeats, and all masked repeats. Reuse under CC BY 4.0.

The percentages in this table come from HERV classification work, RepeatMasker hg38 results, and the ENCODE genome annotation.[1][4][5]

Are these viral sequences still viruses?

No. The viral-origin sequences counted in the 8% figure are inherited parts of the human genome, not active viral infections.

Most HERVs are broken molecular fossils. The 2016 classification paper notes that many HERVs entered primate genomes more than 30 million years ago, that solo LTRs are the most common trace, and that no replication-competent HERVs are known. Some younger HERV-K (HML-2) copies still retain more coding potential, but that is not the same as a modern infectious virus circulating in the genome.[1]

Nor is this "non-human DNA" in the everyday sense. These sequences came from ancient viruses, but once they became fixed and inherited, they became part of human DNA. A person's current viral infections, microbiome DNA, or food DNA are separate from the inherited nuclear genome.

A few ancient viral sequences still matter. The human placental protein syncytin came from a captured retroviral envelope gene, and experimental work has shown that some ERV sequences now help regulate immune-response genes.[7][8]

The short version is: humans are not "8% virus" as organisms. But about 8% of the inherited human genome is recognizably viral in origin.

Sources8
  1. Classification and characterization of human endogenous retroviruses; mosaic forms are common

    Retrovirology · 2016

  2. Long-term reinfection of the human genome by endogenous retroviruses

    Proceedings of the National Academy of Sciences · 2004

  3. Human Genome Assembly GRCh38.p14

    Genome Reference Consortium · July 1, 2026

  4. An integrated encyclopedia of DNA elements in the human genome

    Nature · 2012

  5. Human Homo sapiens Genomic Dataset

    RepeatMasker · July 1, 2026

  6. Endogenous Viral Elements in Animal Genomes

    PLOS Genetics · 2010

  7. Syncytin is a captive retroviral envelope protein involved in human placental morphogenesis

    Nature · 2000

  8. Regulatory evolution of innate immunity through co-option of endogenous retroviruses

    Science · 2016

Cite this article

Broz, M. (2026, July 23). What percentage of the human genome is viral DNA? ProteinIQ. https://proteiniq.io/guides/human-genome-viral-dna

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “What percentage of the human genome is viral DNA?” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
July 1, 2026
Updated
July 23, 2026

Related guides

Browse all guides
DNA and sequencing-read motifs beside coins representing genome sequencing cost.

Genetics · September 24, 2026

How much does it cost to sequence a genome?

A human genome costs about $220 to $450 at university sequencing labs and $399 to $595 as a consumer test. ProteinIQ's survey of US core facilities puts the median lab price at $419 before analysis.

Nuclear DNA magnified to show the base pairs of a double helix.

Genetics · September 19, 2026

How big is the human genome?

Compare human genome size in base pairs, nucleotides, picograms, and gigabytes, with reference assembly totals and verified download sizes.

Ink illustration of a sample tube, DNA, sequence reads, and a clock representing sequencing turnaround.

Genetics · July 30, 2026

How long does whole-genome sequencing take?

Human whole-genome sequencing can generate genome data in hours to about a day. Sample preparation, analysis, interpretation, and reporting can extend the full turnaround to weeks.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing