ProteinIQ
Sign inStart for free
ProteinIQ

Sequence analysis

DNA sequence alignment

Align nucleotide sequences while preserving strand, ambiguity, coordinate, and coding-frame context.

Open workflowCompare alignment types
DNA sequence alignmentWorkflow preview

Inputs

1 required

Methods

3 connected

  1. 01MAFFT · DNA Alignment
  2. 02Clustal Omega
  3. 03StringZilla v5 · DNA Global Scores

Run MAFFT and Clustal Omega DNA alignments plus a separate StringZilla global score matrix.

Use this template

On this page

  • Overview
  • Methods
  • Applications
  • Online workflow
  • Interpretation
  • How it works
  • Inputs & outputs

What is DNA sequence alignment?

DNA sequence alignment is the process of arranging two or more DNA sequences in rows so matching or homologous nucleotide positions appear in shared columns. Computers add gaps where needed to line up the letters and reveal matches, substitutions, insertions, deletions, and conserved regions. The task ranges from two short amplicons to multi-sequence gene families and is distinct from read mapping and whole-genome assembly alignment.

Confirm orientation and sequence provenance before alignment. Reverse-complemented records, mixed genomic and transcript sequences, primer remnants, poor-quality ends, and inconsistent locus boundaries can produce plausible-looking but biologically invalid results. Coding regions may benefit from protein-guided review because arbitrary nucleotide gaps can disrupt codons.

Use MAFFT, Clustal Omega, or MUSCLE5 for homologous nucleotide sets. For two complete genomes or assemblies, use MUMmer4 instead. Report ambiguity handling, aligned length, identity, gaps, and reference coordinates rather than presenting percent identity without the denominator or sequence scope.

When to use DNA sequence alignment

  • Best fit. Genes, amplicons, loci, transcripts, alleles, and nucleotide constructs
  • Required input. DNA FASTA records with consistent strand and comparable boundaries

Benefits of DNA sequence alignment

  • Clear correspondence. Reveals nucleotide substitutions and indels
  • Connected evidence. Supports allele and construct comparison
  • Reusable output. Preserves coding and coordinate context

Primary limitations

  • Method dependence. Strand errors can look plausible
  • Input dependence. Repeats create ambiguous placements
  • Interpretive limit. Read mapping requires different methods

DNA sequence alignment methods

Nucleotide alignment uses match, mismatch, and gap scores, sometimes with models that distinguish transition and transversion patterns. Coding sequences add a reading-frame constraint that ordinary nucleotide alignment does not enforce.

Long assemblies require seed-and-extend genome aligners rather than conventional MSA. Choosing the method by data scale and intended coordinate output prevents a generic alignment from being mistaken for variant calling.

DNA sequence alignment applications

DNA sequence alignment is best suited to genes, amplicons, loci, transcripts, alleles, and nucleotide constructs. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run dna sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Verify records. Confirm source, locus, boundaries, alphabet, and sequence quality.
  2. Normalize orientation. Orient homologs consistently and remove primers or unrelated flanks when appropriate.
  3. Choose nucleotide method. Choose pairwise, multiple, global, local, or genome-scale behavior.
  4. Run and inspect. Generate alignments and inspect gaps, ambiguity, repeats, and coding frames.
  5. Report coordinates. Report identity denominator, coordinate convention, settings, and exclusions.

How to interpret dna sequence alignment results

State whether identity excludes gaps and ambiguous bases. Review substitutions in codon context when the sequence encodes protein, and distinguish synonymous from amino-acid-changing differences downstream.

Do not infer a variant from an alignment alone without sequence-quality and reference checks. Alignment exposes candidate differences; provenance and validation establish whether they are real.

How dna sequence alignment works

Run MAFFT and Clustal Omega DNA alignments plus a separate StringZilla global score matrix.

  1. Verify records. Confirm source, locus, boundaries, alphabet, and sequence quality.
  2. Normalize orientation. Orient homologs consistently and remove primers or unrelated flanks when appropriate.
  3. Choose nucleotide method. Choose pairwise, multiple, global, local, or genome-scale behavior.
  4. Run and inspect. Generate alignments and inspect gaps, ambiguity, repeats, and coding frames.
  5. Report coordinates. Report identity denominator, coordinate convention, settings, and exclusions.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF DNA FASTA records with consistent orientation, boundaries, and identifiers.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Aligned nucleotide FASTA, global score matrices, gap patterns, and review files.

Tools for DNA sequence alignment

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

MAFFT

Create configurable protein, DNA, or RNA alignments

Clustal Omega

Create scalable protein or nucleotide multiple-sequence alignments

StringZilla v5

Calculate pairwise global, local, or edit-distance score matrices

MUSCLE5

Generate conventional or ensemble multiple-sequence alignments

MUMmer4

Align genome assemblies and report coordinates and variants

MMseqs2

Search and cluster large protein or nucleotide sequence sets

HMMER

Find homologs with profile hidden Markov models

FastTree

Estimate trees from large sequence alignments

IQ-TREE

Infer maximum-likelihood phylogenies from alignments

RAxML-NG

Run maximum-likelihood phylogenetic analysis

RNAalifold

Predict consensus RNA structure from an RNA alignment

IgBLAST

Annotate immunoglobulin and T-cell receptor rearrangements

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Pairwise sequence alignment

Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.

Multiple sequence alignment

Aligns three or more homologous sequences to identify shared positions, insertions, deletions, and conserved regions.

Global sequence alignment

Compares sequences end to end, including terminal differences and gaps across their full lengths.

Protein sequence alignment

Aligns amino-acid sequences using substitution-aware methods suited to protein evolution and function.

Local sequence alignment

Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.

Whole genome alignment

Maps large homologous regions between genome assemblies and reports coordinates, rearrangements, and sequence differences.

Structure-based sequence alignment

Uses three-dimensional correspondence to align residues whose sequence similarity alone may be weak.

Frequently asked questions

DNA FASTA records with consistent orientation, boundaries, and identifiers.

Aligned nucleotide FASTA, global score matrices, gap patterns, and review files.

Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.

No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.

Preserve strand, reference build or source, boundaries, ambiguity rules, scoring settings, and coordinate conventions.

A complete dna sequence alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.

The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.

ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured dna sequence alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow
ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing