ProteinIQ
Sign inStart for free
ProteinIQ

Sequence analysis

Global sequence alignment

Compare complete sequences end to end and make terminal gaps, full-length coverage, and score definitions explicit.

Open workflowCompare alignment types
Global sequence alignmentWorkflow preview

Inputs

1 required

Methods

2 connected

  1. 01StringZilla v5 · Needleman–Wunsch Score
  2. 02MAFFT · G-INS-i Alignment

Compare StringZilla Needleman–Wunsch global scores with a separate MAFFT G-INS-i alignment.

Use this template

On this page

  • Overview
  • Methods
  • Applications
  • Online workflow
  • Interpretation
  • How it works
  • Inputs & outputs

What is global sequence alignment?

Global sequence alignment is the process of arranging sequences from beginning to end so their complete lengths are compared. Classical Needleman–Wunsch dynamic programming finds an optimal end-to-end path under specified substitution and gap scores. Global multiple-alignment strategies extend the same full-length assumption to a sequence set. The approach is most defensible when sequences have comparable boundaries and domain architecture.

Choose global alignment for full-length homologs, alleles, orthologs, or engineered constructs where terminal and internal differences all matter. If sequences share only one domain or motif, forcing unrelated flanks into an end-to-end comparison can create long gaps and misleading identity values.

Interpret a global score only with its scoring system and sequence lengths. Protein substitution matrices, nucleotide match values, gap opening, gap extension, and ambiguous-symbol handling can change the optimal path. Review aligned strings alongside scores because a matrix alone does not reveal which residues were paired.

When to use global sequence alignment

  • Best fit. Comparable full-length homologs and complete construct comparisons
  • Required input. Sequences with matched boundaries and broadly similar architecture

Benefits of global sequence alignment

  • Clear correspondence. Uses every residue in the comparison
  • Connected evidence. Highlights terminal and internal differences
  • Reusable output. Fits similar full-length homologs

Primary limitations

  • Method dependence. Poor fit for fragments or domain sharing
  • Input dependence. Length differences dominate some results
  • Interpretive limit. Scores are parameter-specific

Global sequence alignment methods

Needleman–Wunsch fills a dynamic-programming matrix and traces an optimal path from one end to the other. Affine gaps distinguish the cost of opening a gap from extending it.

MAFFT G-INS-i is a global-homology multiple-alignment strategy, not a synonym for a pairwise Needleman–Wunsch traceback. Use the former for aligned sequence sets and the latter when the exact pairwise scoring model is required.

Global sequence alignment applications

Global sequence alignment is best suited to comparable full-length homologs and complete construct comparisons. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run global sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Confirm full-length scope. Confirm that full-length correspondence answers the biological question.
  2. Normalize inputs. Check sequence boundaries, orientation, molecule type, and ambiguous symbols.
  3. Set end-to-end method. Choose and record substitution and affine-gap parameters.
  4. Run score and alignment. Calculate global scores and generate an inspectable alignment.
  5. Review terminal effects. Review coverage, termini, internal gaps, identity, and functional positions.

How to interpret global sequence alignment results

Report score, aligned length, identities, substitutions, gap opens, and coverage. Normalize cautiously because different normalization formulas answer different questions.

Long terminal gaps often indicate unmatched boundaries rather than biological insertions. Revisit sequence extraction before drawing evolutionary or functional conclusions.

How global sequence alignment works

Compare StringZilla Needleman–Wunsch global scores with a separate MAFFT G-INS-i alignment.

  1. Confirm full-length scope. Confirm that full-length correspondence answers the biological question.
  2. Normalize inputs. Check sequence boundaries, orientation, molecule type, and ambiguous symbols.
  3. Set end-to-end method. Choose and record substitution and affine-gap parameters.
  4. Run score and alignment. Calculate global scores and generate an inspectable alignment.
  5. Review terminal effects. Review coverage, termini, internal gaps, identity, and functional positions.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Full-length protein or nucleotide FASTA sequences with compatible boundaries.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Global score matrices and separately generated end-to-end aligned sequences.

Tools for global sequence alignment

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

StringZilla v5

Calculate pairwise global, local, or edit-distance score matrices

MAFFT

Create configurable protein, DNA, or RNA alignments

Clustal Omega

Create scalable protein or nucleotide multiple-sequence alignments

MUSCLE5

Generate conventional or ensemble multiple-sequence alignments

PDB to FASTA converter

Extract sequences from structures before comparison

HMMER

Find homologs with profile hidden Markov models

MMseqs2

Search and cluster large protein or nucleotide sequence sets

USAlign

Align macromolecular structures and derive residue correspondence

FoldSeek

Find and compare structurally similar proteins

IQ-TREE

Infer maximum-likelihood phylogenies from alignments

FastTree

Estimate trees from large sequence alignments

RAxML-NG

Run maximum-likelihood phylogenetic analysis

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Pairwise sequence alignment

Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.

Multiple sequence alignment

Aligns three or more homologous sequences to identify shared positions, insertions, deletions, and conserved regions.

Protein sequence alignment

Aligns amino-acid sequences using substitution-aware methods suited to protein evolution and function.

DNA sequence alignment

Aligns nucleotide sequences to compare homologous genes, amplicons, loci, transcripts, or constructs.

Local sequence alignment

Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.

Whole genome alignment

Maps large homologous regions between genome assemblies and reports coordinates, rearrangements, and sequence differences.

Structure-based sequence alignment

Uses three-dimensional correspondence to align residues whose sequence similarity alone may be weak.

Frequently asked questions

Full-length protein or nucleotide FASTA sequences with compatible boundaries.

Global score matrices and separately generated end-to-end aligned sequences.

Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.

No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.

Record the score matrix, gap model, sequence boundaries, alignment path source, and all normalization choices.

A complete global sequence alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.

The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.

ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured global sequence alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow
ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing