ProteinIQ
Sign inStart for free
ProteinIQ
Statistics/Aug '26/5 min read

How accurate is DNA sequencing?

Matic Broz

Matic BrozComputational chemist

Modern DNA sequencing is usually about 99% to more than 99.9% accurate per base under defined conditions. The percentage depends on the platform, chemistry, software, sample, and whether it describes a raw read or a consensus built from repeated observations.

A bare claim such as “99.9% accurate” is incomplete. It says nothing about which parts of a genome were missed, whether an assembly joined them correctly, or how well a pipeline identified variants.

How accurate is DNA sequencing?

Modern sequencing platforms commonly report high-quality base calls between Q20 and Q40, equal to expected error probabilities from 1 in 100 to 1 in 10,000 bases.[1][2]

Current published figures are not direct competitors unless the metric and test conditions match:

Platform data productPublished accuracy figureConfiguration, sample, and metric
Illumina NextSeq 2000 short readsAt least 90% of bases above Q30; at least 80% above Q40P4 XLEAP-SBS reagents with RTA 4; vendor system specification; test sample is not named in the public note
Oxford Nanopore simplex and duplex readsSimplex up to 99%; duplex about Q30, or 99.9%R10.4.1 flow cell, Kit 14, SUP basecalling, and Dorado model dna_r10.4.1_e8.2_400bps_sup@v4.1.0; vendor example; sample is not named
PacBio HiFi readsUp to 99.95% average read accuracyCurrent vendor comparison; chemistry, analysis version, and benchmark sample are not stated on the page
Sanger capillary reads99.99%Vendor educational claim; instrument, chemistry, sample, read position, and trimming rule are not stated

The figures above come from first-party technical or educational pages, so they describe each vendor's own stated performance rather than one independent, controlled benchmark.[3][4][5][7]

Phred scores put expected base-call errors on a logarithmic scale: Q20 means 1% error probability, Q30 means 0.1%, Q40 means 0.01%, and Q50 means 0.001%.[1][2]

Expected base-call error probability falls tenfold with each 10-point increase in Phred quality score

These are probabilities assigned to individual base calls or summaries of those probabilities. Q30 does not mean that every 1,000-base read contains exactly one error, and the percentage of Q30 bases is not the same measurement as whole-read identity.

How accurate is Illumina sequencing?

Illumina's current NextSeq 2000 specification reports at least 90% of bases above Q30 and at least 80% above Q40 with P4 XLEAP-SBS reagents and RTA 4.[3]

Q30 corresponds to 99.9% inferred accuracy for that base call, while Q40 corresponds to 99.99%. The stated percentages describe the share of bases crossing each threshold, not the accuracy of every read or every position in a genome.

The accuracy of next generation sequencing platforms also changes with the library, read length, cluster loading, sequence context, and run quality. A high Q30 fraction cannot recover a region that received no usable reads. More coverage can reduce random uncertainty, but it adds data and sequencing cost.

Illumina stores a quality character beside every called base in FASTQ data. FASTQ to FASTA conversion removes those quality scores, so the resulting FASTA file cannot preserve the original per-base confidence.

How accurate are Nanopore and PacBio long reads?

Oxford Nanopore's Kit 14 guide reports simplex reads up to 99% accurate and duplex reads around Q30 when R10.4.1 flow cells are paired with SUP basecalling.[4]

The guide names Dorado model dna_r10.4.1_e8.2_400bps_sup@v4.1.0 for its duplex workflow. Duplex basecalling combines the template and complement strands, giving the software two observations of the same molecule. Duplex reads make up only part of a run, so their Q30 figure should not be applied to every Nanopore read.

Nanopore accuracy is especially sensitive to chemistry and basecaller version. The 2020 PrecisionFDA benchmark used older R9.4.1 chemistry and Guppy 3.6.0, so its results should not be presented as the accuracy of current R10.4.1 and Dorado data.[8]

PacBio's current product page claims up to 99.95% average accuracy for HiFi reads but does not name the chemistry, analysis version, or sample behind that comparison.[5] A peer-reviewed 2019 benchmark provides clearer context: circular consensus sequencing of the human HG002 genome produced reads averaging 13.5 kb and 99.8% accuracy with the then-current Sequel II and CCS workflow.[6]

HiFi derives a consensus by reading a circularized molecule repeatedly. That makes it a consensus long read, rather than a single observation that can be compared directly with a Nanopore simplex read.

How accurate is Sanger sequencing?

Thermo Fisher describes Sanger sequencing as 99.99% accurate, but that first-party figure is not a guarantee for every base in a capillary trace.[7]

Sanger quality varies along the trace. The ends and unresolved peaks can carry much higher error probabilities than the clean middle region, and mixed templates can produce overlapping peaks. The original Phred studies assigned a calibrated error probability to each base rather than treating a trace as one fixed percentage.[1]

Sanger sequencing accuracy therefore needs a rule such as a minimum Q score, a trimmed read interval, agreement between forward and reverse reads, or a consensus sequence. Calling the method “99.99% accurate” without one of those qualifiers overstates what a single trace establishes.

Does read accuracy predict assembly and variant accuracy?

Read accuracy does not predict assembly or variant accuracy by itself. Raw-read accuracy influences later analysis, but the finished data products have their own denominators and benchmarks.

An assembly consensus QV estimates base errors in the assembled sequence. Completeness, contiguity, misjoins, and phasing remain separate properties. Merqury, for example, reports consensus QV and k-mer completeness separately because an assembly can spell retained sequence accurately while still omitting or misplacing other sequence.[9] Researchers can compare assemblies directly, but a high read Q score alone cannot establish assembly correctness.

Variant callers are evaluated with precision, recall, and their harmonic mean, the F1 score, against a truth set. In the 2020 PrecisionFDA challenge, teams analyzed about 35-fold Illumina, 35-fold PacBio HiFi, and 50-fold Nanopore data for Genome in a Bottle samples. The best single-platform submissions reached F1 scores of 0.998 for PacBio HiFi and 0.997 for Illumina across all benchmark regions, but the corresponding difficult-to-map results were 0.993 and 0.969.[8]

Those figures belong to specific samples, coverages, callers, references, and benchmark regions. They measure small-variant calls from complete pipelines, not raw base accuracy and not the fraction of the human genome that any platform can resolve.

Sources▼
  1. Base-Calling of Automated Sequencer Traces Using Phred. II. Error Probabilities Genome Research · 1998. https://genome.cshlp.org/content/8/3/186.full
  2. Sequencing Quality Scores Illumina · August 10, 2026. https://www.illumina.com/science/technology/next-generation-sequencing/plan-experiments/quality-scores.html
  3. Superior performance with the NextSeq 2000 System Illumina · August 10, 2026. https://support.illumina.com/content/dam/illumina/gcs/assembled-assets/marketing-literature/nextseq-2000-aviti-performance-tech-note-m-gl-03083/nextseq-2000-aviti-performance-technical-note-m-gl-03083.pdf
  4. Kit 14 sequencing and duplex basecalling Oxford Nanopore Technologies · August 10, 2026. https://nanoporetech.com/es/document/kit-14-device-and-informatics/
  5. Long-read sequencing: Benefits and HiFi accuracy PacBio · August 10, 2026. https://www.pacb.com/technology/long-read-sequencing/
  6. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome Nature Biotechnology · 2019. https://www.nature.com/articles/s41587-019-0217-9
  7. What is Sanger sequencing? Thermo Fisher Scientific · August 10, 2026. https://www.thermofisher.com/us/en/home/life-science/sequencing/sequencing-learning-center/capillary-electrophoresis-information/what-is-sanger-sequencing.html
  8. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions Cell Genomics · 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC9205427/
  9. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies Genome Biology · 2020. https://link.springer.com/article/10.1186/s13059-020-02134-9
Published
August 10, 2026
Last updated
August 10, 2026

Table of contents

Cite this article

Broz, M. (2026, August 10). How accurate is DNA sequencing? ProteinIQ. https://proteiniq.io/guides/dna-sequencing-accuracy

Matic Broz, PhD

Matic Broz, PhD

Founder and computational chemist, ProteinIQ

Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.

Related guides

Statistics

JUL '26

How much DNA is needed for sequencing?

Most DNA sequencing workflows start with 1 ng to 1 µg of DNA. Illumina DNA Prep recommends 100–500 ng for human DNA, while Nanopore ligation libraries use 1 µg for fragments longer than 10 kb.

Matic Broz Computational chemist

Statistics

JUL '26

How long does whole-genome sequencing take?

Human whole-genome sequencing can generate genome data in hours to about a day. Sample preparation, analysis, interpretation, and reporting can extend the full turnaround to weeks.

Matic Broz Computational chemist

Statistics

JUL '26

How big is the human genome?

The human genome has about 3.1 billion base pairs per haploid copy. A typical diploid cell contains about 6.3 billion base pairs and 6.4 to 6.5 picograms of nuclear DNA.

Matic Broz Computational chemist

ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing