Use case

Global sequence alignment

Compare complete sequences end to end and make terminal gaps, full-length coverage, and score definitions explicit.

Global sequence alignmentRead-only preview

Inputs

1 required

Methods

2 connected

  1. 01StringZilla v5 · Needleman–Wunsch Score
  2. 02MAFFT · G-INS-i Alignment

Compare StringZilla Needleman–Wunsch global scores with a separate MAFFT G-INS-i alignment.

Use this template

What is global sequence alignment?

Global sequence alignment is the process of arranging sequences from beginning to end so their complete lengths are compared. Classical Needleman–Wunsch dynamic programming finds an optimal end-to-end path under specified substitution and gap scores. Global multiple-alignment strategies extend the same full-length assumption to a sequence set. The approach is most defensible when sequences have comparable boundaries and domain architecture.

Choose global alignment for full-length homologs, alleles, orthologs, or engineered constructs where terminal and internal differences all matter. If sequences share only one domain or motif, forcing unrelated flanks into an end-to-end comparison can create long gaps and misleading identity values.

Interpret a global score only with its scoring system and sequence lengths. Protein substitution matrices, nucleotide match values, gap opening, gap extension, and ambiguous-symbol handling can change the optimal path. Review aligned strings alongside scores because a matrix alone does not reveal which residues were paired.

When to use global sequence alignment

  • Best fit. Comparable full-length homologs and complete construct comparisons
  • Required input. Sequences with matched boundaries and broadly similar architecture

Benefits of global sequence alignment

  • Clear correspondence. Uses every residue in the comparison
  • Connected evidence. Highlights terminal and internal differences
  • Reusable output. Fits similar full-length homologs

Primary limitations

  • Method dependence. Poor fit for fragments or domain sharing
  • Input dependence. Length differences dominate some results
  • Interpretive limit. Scores are parameter-specific

Global sequence alignment methods

Needleman–Wunsch fills a dynamic-programming matrix and traces an optimal path from one end to the other. Affine gaps distinguish the cost of opening a gap from extending it.

MAFFT G-INS-i is a global-homology multiple-alignment strategy, not a synonym for a pairwise Needleman–Wunsch traceback. Use the former for aligned sequence sets and the latter when the exact pairwise scoring model is required.

Global sequence alignment applications

Global sequence alignment is best suited to comparable full-length homologs and complete construct comparisons. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run global sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Confirm full-length scope. Confirm that full-length correspondence answers the biological question.
  2. Normalize inputs. Check sequence boundaries, orientation, molecule type, and ambiguous symbols.
  3. Set end-to-end method. Choose and record substitution and affine-gap parameters.
  4. Run score and alignment. Calculate global scores and generate an inspectable alignment.
  5. Review terminal effects. Review coverage, termini, internal gaps, identity, and functional positions.

How to interpret global sequence alignment results

Report score, aligned length, identities, substitutions, gap opens, and coverage. Normalize cautiously because different normalization formulas answer different questions.

Long terminal gaps often indicate unmatched boundaries rather than biological insertions. Revisit sequence extraction before drawing evolutionary or functional conclusions.

How global sequence alignment works

Compare StringZilla Needleman–Wunsch global scores with a separate MAFFT G-INS-i alignment.

  1. Confirm full-length scope. Confirm that full-length correspondence answers the biological question.
  2. Normalize inputs. Check sequence boundaries, orientation, molecule type, and ambiguous symbols.
  3. Set end-to-end method. Choose and record substitution and affine-gap parameters.
  4. Run score and alignment. Calculate global scores and generate an inspectable alignment.
  5. Review terminal effects. Review coverage, termini, internal gaps, identity, and functional positions.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Full-length protein or nucleotide FASTA sequences with compatible boundaries.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Global score matrices and separately generated end-to-end aligned sequences.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow