ProteinIQ
Get a demoSign inStart for free
ProteinIQ
Genetics

How accurate is DNA sequencing?

August 10, 2026·Matic Broz, PhD
A magnifying glass examines DNA alongside overlapping sequence fragments.

Modern DNA sequencing is usually about 99% to more than 99.9% accurate per base under defined conditions. The percentage depends on the platform, chemistry, software, sample, and whether it describes a raw read or a consensus built from repeated observations.

A bare claim such as “99.9% accurate” is incomplete. It says nothing about which parts of a genome were missed, whether an assembly joined them correctly, or how well a pipeline identified variants.

How accurate is DNA sequencing?

Modern sequencing platforms commonly report high-quality base calls between Q20 and Q40, equal to expected error probabilities from 1 in 100 to 1 in 10,000 bases.[1][2]

Current published figures are not direct competitors unless the metric and test conditions match:

Platform data productPublished accuracy figureConfiguration, sample, and metric
Illumina NextSeq 2000 short readsAt least 90% of bases above Q30; at least 80% above Q40P4 XLEAP-SBS reagents with RTA 4; vendor system specification; test sample is not named in the public note
Oxford Nanopore simplex and duplex readsSimplex up to 99%; duplex about Q30, or 99.9%R10.4.1 flow cell, Kit 14, SUP basecalling, and Dorado model dna_r10.4.1_e8.2_400bps_sup@v4.1.0; vendor example; sample is not named
PacBio HiFi readsUp to 99.95% average read accuracyCurrent vendor comparison; chemistry, analysis version, and benchmark sample are not stated on the page
Sanger capillary reads99.99%Vendor educational claim; instrument, chemistry, sample, read position, and trimming rule are not stated

The figures above come from first-party technical or educational pages, so they describe each vendor's own stated performance rather than one independent, controlled benchmark.[3][4][5][7]

Phred scores put expected base-call errors on a logarithmic scale: Q20 means 1% error probability, Q30 means 0.1%, Q40 means 0.01%, and Q50 means 0.001%.[1][2]

Expected base-call error probability falls tenfold with each 10-point increase in Phred quality score. Reuse under CC BY 4.0.

These are probabilities assigned to individual base calls or summaries of those probabilities. Q30 does not mean that every 1,000-base read contains exactly one error, and the percentage of Q30 bases is not the same measurement as whole-read identity.

How accurate is Illumina sequencing?

Illumina's current NextSeq 2000 specification reports at least 90% of bases above Q30 and at least 80% above Q40 with P4 XLEAP-SBS reagents and RTA 4.[3]

Q30 corresponds to 99.9% inferred accuracy for that base call, while Q40 corresponds to 99.99%. The stated percentages describe the share of bases crossing each threshold, not the accuracy of every read or every position in a genome.

The accuracy of next generation sequencing platforms also changes with the library, read length, cluster loading, sequence context, and run quality. A high Q30 fraction cannot recover a region that received no usable reads. More coverage can reduce random uncertainty, but it adds data and sequencing cost.

Illumina stores a quality character beside every called base in FASTQ data. FASTQ to FASTA conversion removes those quality scores, so the resulting FASTA file cannot preserve the original per-base confidence.

How accurate are Nanopore and PacBio long reads?

Oxford Nanopore's Kit 14 guide reports simplex reads up to 99% accurate and duplex reads around Q30 when R10.4.1 flow cells are paired with SUP basecalling.[4]

The guide names Dorado model dna_r10.4.1_e8.2_400bps_sup@v4.1.0 for its duplex workflow. Duplex basecalling combines the template and complement strands, giving the software two observations of the same molecule. Duplex reads make up only part of a run, so their Q30 figure should not be applied to every Nanopore read.

Nanopore accuracy is especially sensitive to chemistry and basecaller version. The 2020 PrecisionFDA benchmark used older R9.4.1 chemistry and Guppy 3.6.0, so its results should not be presented as the accuracy of current R10.4.1 and Dorado data.[8]

PacBio's current product page claims up to 99.95% average accuracy for HiFi reads but does not name the chemistry, analysis version, or sample behind that comparison.[5] A peer-reviewed 2019 benchmark provides clearer context: circular consensus sequencing of the human HG002 genome produced reads averaging 13.5 kb and 99.8% accuracy with the then-current Sequel II and CCS workflow.[6]

HiFi derives a consensus by reading a circularized molecule repeatedly. That makes it a consensus long read, rather than a single observation that can be compared directly with a Nanopore simplex read.

How accurate is Sanger sequencing?

Thermo Fisher describes Sanger sequencing as 99.99% accurate, but that first-party figure is not a guarantee for every base in a capillary trace.[7]

Sanger quality varies along the trace. The ends and unresolved peaks can carry much higher error probabilities than the clean middle region, and mixed templates can produce overlapping peaks. The original Phred studies assigned a calibrated error probability to each base rather than treating a trace as one fixed percentage.[1]

Sanger sequencing accuracy therefore needs a rule such as a minimum Q score, a trimmed read interval, agreement between forward and reverse reads, or a consensus sequence. Calling the method “99.99% accurate” without one of those qualifiers overstates what a single trace establishes.

Does read accuracy predict assembly and variant accuracy?

Read accuracy does not predict assembly or variant accuracy by itself. Raw-read accuracy influences later analysis, but the finished data products have their own denominators and benchmarks.

An assembly consensus QV estimates base errors in the assembled sequence. Completeness, contiguity, misjoins, and phasing remain separate properties. Merqury, for example, reports consensus QV and k-mer completeness separately because an assembly can spell retained sequence accurately while still omitting or misplacing other sequence.[9] Researchers can compare assemblies directly, but a high read Q score alone cannot establish assembly correctness.

Variant callers are evaluated with precision, recall, and their harmonic mean, the F1 score, against a truth set. In the 2020 PrecisionFDA challenge, teams analyzed about 35-fold Illumina, 35-fold PacBio HiFi, and 50-fold Nanopore data for Genome in a Bottle samples. The best single-platform submissions reached F1 scores of 0.998 for PacBio HiFi and 0.997 for Illumina across all benchmark regions, but the corresponding difficult-to-map results were 0.993 and 0.969.[8]

Those figures belong to specific samples, coverages, callers, references, and benchmark regions. They measure small-variant calls from complete pipelines, not raw base accuracy and not the fraction of the human genome that any platform can resolve.

Sources9
  1. Base-Calling of Automated Sequencer Traces Using Phred. II. Error Probabilities

    Genome Research · 1998

  2. Sequencing Quality Scores

    Illumina · August 10, 2026

  3. Superior performance with the NextSeq 2000 System

    Illumina · August 10, 2026

  4. Kit 14 sequencing and duplex basecalling

    Oxford Nanopore Technologies · August 10, 2026

  5. Long-read sequencing: Benefits and HiFi accuracy

    PacBio · August 10, 2026

  6. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome

    Nature Biotechnology · 2019

  7. What is Sanger sequencing?

    Thermo Fisher Scientific · August 10, 2026

  8. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions

    Cell Genomics · 2022

  9. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies

    Genome Biology · 2020

Cite this article

Broz, M. (2026, August 10). How accurate is DNA sequencing? ProteinIQ. https://proteiniq.io/guides/dna-sequencing-accuracy

Reuse the chartsCC BY 4.0

You can use the charts in this article in your own articles, slides and teaching materials, including commercial work, under the CC BY 4.0 license. Credit ProteinIQ and link to this page. The license covers the charts only, not the article text or illustrations.

Credit

Chart: “How accurate is DNA sequencing?” by ProteinIQ, CC BY 4.0

About the author

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

  • LinkedIn
  • Google Scholar
  • ORCID
Published
August 10, 2026

Related guides

Browse all guides
DNA and sequencing-read motifs beside coins representing genome sequencing cost.

Genetics · September 24, 2026

How much does it cost to sequence a genome?

A human genome costs about $220 to $450 at university sequencing labs and $399 to $595 as a consumer test. ProteinIQ's survey of US core facilities puts the median lab price at $419 before analysis.

Pipette and sample tube beside a DNA helix for sequencing preparation.

Genetics · July 28, 2026

How much DNA is needed for sequencing?

Most DNA sequencing workflows start with 1 ng to 1 µg of DNA. Illumina DNA Prep recommends 100–500 ng for human DNA, while Nanopore ligation libraries use 1 µg for fragments longer than 10 kb.

Nuclear DNA magnified to show the base pairs of a double helix.

Genetics · September 19, 2026

How big is the human genome?

Compare human genome size in base pairs, nucleotides, picograms, and gigabytes, with reference assembly totals and verified download sizes.

ProteinIQ

Published bioinformatics tools, ready to run in the browser.

Platform

  • Bioinformatics tools
  • Workflows
  • Batches
  • AI Assistant
  • PDB viewer

Developers

  • Examples
  • API
  • Python SDK
  • MCP server

Popular tools

  • Boltz-2
  • AlphaFold 2
  • ESMFold
  • AutoDock Vina
  • RFdiffusion
  • ProteinMPNN
  • All tools

Teams

  • For academia
  • For enterprise

Research areas

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering

Use cases

  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • Protein design
  • Molecular dynamics simulation
  • All use cases

Resources

  • Documentation
  • Guides
  • Datasets
  • Blog
  • Customers
  • Changelog
  • Sitemap

Company

  • About
  • Careers
  • Contact
  • Pricing
  • Author

Trust and legal

  • Security
  • Trust center
  • Terms
  • Privacy policy
  • All legal documents

© 2026 ProteinIQ

  • Pricing