ProteinIQ
Sign inStart for free
ProteinIQ

Sequence analysis

Whole genome alignment

Compare genome assemblies with coordinate-aware alignment, dot plots, variant tables, and retained delta files.

Open workflowCompare alignment types
Whole genome alignmentWorkflow preview

Inputs

2 required

Methods

1 connected

  1. 01MUMmer4 · NUCmer Genome Alignment

MUMmer4 NUCmer aligns reference and query assemblies and returns coordinates, delta files, variants, statistics, and dot plots.

Use this template

On this page

  • Overview
  • Methods
  • Applications
  • Online workflow
  • Interpretation
  • How it works
  • Inputs & outputs

What is whole genome alignment?

Whole genome alignment is the process of arranging large genome assemblies or contig sets against one another to identify corresponding blocks and their coordinates. It can reveal collinearity, rearrangements, substitutions, insertions, and deletions. MUMmer4 NUCmer uses exact matches as anchors and extends them into alignments, producing assembly-to-assembly evidence rather than a substitute for raw-read mapping or independently validated variant calls.

Use whole-genome alignment for comparing strains, assembly versions, haplotypes, related species, or a draft assembly against a reference. Reference choice, assembly quality, repeat content, ploidy, and evolutionary distance determine how coordinates and apparent rearrangements should be interpreted.

Review dot plots and coordinate tables before focusing on variants. Duplications, inversions, translocations, unplaced contigs, and repeat-driven multi-mapping can create several valid correspondences. Preserve delta files, filters, reference names, and coordinate conventions so every reported difference can be traced back to the alignment.

When to use whole genome alignment

  • Best fit. Assembly comparison, synteny, rearrangements, polishing, and candidate differences
  • Required input. Reference and query genome or contig FASTA files with assembly metadata

Benefits of whole genome alignment

  • Clear correspondence. Scales to genome assemblies
  • Connected evidence. Reports coordinate-aware differences
  • Reusable output. Visualizes synteny and rearrangements

Primary limitations

  • Method dependence. Assembly errors mimic variation
  • Input dependence. Repeats create ambiguous matches
  • Interpretive limit. Raw-read support is not included

Whole genome alignment methods

NUCmer identifies maximal exact matches between nucleotide sequences and clusters consistent anchors before extension. Match-mode and filtering settings change sensitivity and the number of repeated correspondences retained.

Dot plots summarize large-scale geometry, while coordinate and delta files preserve detailed alignments. SNP and indel reports should be interpreted only after repetitive and multiply aligned regions are understood.

Whole genome alignment applications

Whole genome alignment is best suited to assembly comparison, synteny, rearrangements, polishing, and candidate differences. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run whole genome alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Select assemblies. Choose a biologically appropriate reference and document assembly versions.
  2. Check quality. Review contiguity, contamination, haplotype status, and repeat content.
  3. Run NUCmer. Run MUMmer4 with explicit match mode, strand, and filtering settings.
  4. Review genome geometry. Inspect dot plots, coverage, coordinates, duplications, and rearrangements.
  5. Validate differences. Validate candidate variants or breakpoints with reads and orthogonal evidence.

How to interpret whole genome alignment results

A clean diagonal suggests broadly collinear assemblies; reverse diagonals indicate inversions, and off-diagonal blocks may represent translocations, duplications, or assembly placement differences.

Confirm major events using assembly graphs, raw-read mappings, or an independent method. Coordinate conventions and reference direction must remain explicit in every exported table.

How whole genome alignment works

MUMmer4 NUCmer aligns reference and query assemblies and returns coordinates, delta files, variants, statistics, and dot plots.

  1. Select assemblies. Choose a biologically appropriate reference and document assembly versions.
  2. Check quality. Review contiguity, contamination, haplotype status, and repeat content.
  3. Run NUCmer. Run MUMmer4 with explicit match mode, strand, and filtering settings.
  4. Review genome geometry. Inspect dot plots, coverage, coordinates, duplications, and rearrangements.
  5. Validate differences. Validate candidate variants or breakpoints with reads and orthogonal evidence.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Reference and query genome or contig assemblies in FASTA format.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Delta and coordinate files, SNP and indel tables, alignment statistics, and dot plots.

Tools for whole genome alignment

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

MUMmer4

Align genome assemblies and report coordinates and variants

MAFFT

Create configurable protein, DNA, or RNA alignments

Clustal Omega

Create scalable protein or nucleotide multiple-sequence alignments

StringZilla v5

Calculate pairwise global, local, or edit-distance score matrices

MMseqs2

Search and cluster large protein or nucleotide sequence sets

HMMER

Find homologs with profile hidden Markov models

MUSCLE5

Generate conventional or ensemble multiple-sequence alignments

FastTree

Estimate trees from large sequence alignments

IQ-TREE

Infer maximum-likelihood phylogenies from alignments

RAxML-NG

Run maximum-likelihood phylogenetic analysis

RNAalifold

Predict consensus RNA structure from an RNA alignment

USAlign

Align macromolecular structures and derive residue correspondence

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Pairwise sequence alignment

Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.

Multiple sequence alignment

Aligns three or more homologous sequences to identify shared positions, insertions, deletions, and conserved regions.

Global sequence alignment

Compares sequences end to end, including terminal differences and gaps across their full lengths.

Protein sequence alignment

Aligns amino-acid sequences using substitution-aware methods suited to protein evolution and function.

DNA sequence alignment

Aligns nucleotide sequences to compare homologous genes, amplicons, loci, transcripts, or constructs.

Local sequence alignment

Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.

Structure-based sequence alignment

Uses three-dimensional correspondence to align residues whose sequence similarity alone may be weak.

Frequently asked questions

Reference and query genome or contig assemblies in FASTA format.

Delta and coordinate files, SNP and indel tables, alignment statistics, and dot plots.

Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.

No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.

Retain assembly accessions and versions, MUMmer4 settings, delta files, filters, coordinate conventions, and orthogonal validation.

A complete whole genome alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.

The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.

ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured whole genome alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow
ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing