Use case
Multiple sequence alignment
Align homologous sequence sets with multiple methods, review conserved columns and gaps, and keep method uncertainty visible.
Inputs
1 required
Methods
3 connected
- 01Clustal Omega
- 02MAFFT · L-INS-i
- 03MUSCLE5
Run Clustal Omega, MAFFT L-INS-i, and MUSCLE5 from one FASTA set and compare their native alignments.
Use this templateWhat is multiple sequence alignment?
Multiple sequence alignment is the process of arranging three or more homologous protein, DNA, or RNA sequences in rows so corresponding positions appear in shared columns. Each column is a hypothesis of evolutionary or structural correspondence. The alignment supports motif analysis, profile construction, phylogenetics, consensus calculation, and annotation transfer, but it is not direct evidence that every column is correct.
The most important input decision is which sequences belong together. Paralog mixtures, fragments, nonhomologous domains, duplicated records, and extreme length differences can distort guide trees and gap placement. Curate the set before tuning an algorithm, and consider aligning domains separately when architectures differ.
Compare methods or parameterizations when downstream conclusions depend on uncertain regions. Agreement among Clustal Omega, MAFFT, and MUSCLE5 can increase confidence in stable blocks, while disagreement identifies columns that should be masked, structurally checked, or excluded from tree inference and residue-level claims.
When to use multiple sequence alignment
- Best fit. Protein families, conserved motifs, consensus sequences, profiles, and phylogenetic preparation
- Required input. Three or more curated homologous sequences of one molecular type
Benefits of multiple sequence alignment
- Clear correspondence. Reveals conserved positions across a family
- Connected evidence. Supports profiles and phylogenies
- Reusable output. Allows method sensitivity checks
Primary limitations
- Method dependence. Column homology is inferred
- Input dependence. Input composition strongly affects results
- Interpretive limit. Large divergent sets remain difficult
Multiple sequence alignment methods
Progressive methods build an initial guide tree and add sequences or profiles in stages. Iterative refinement can revisit earlier decisions, while consistency and profile techniques use additional evidence to improve difficult alignments.
MAFFT modes trade speed against accuracy and differ in how they treat global homology and long gaps. MUSCLE5 can generate alignment ensembles, and Clustal Omega scales through profile-based progressive alignment.
Multiple sequence alignment applications
Multiple sequence alignment is best suited to protein families, conserved motifs, consensus sequences, profiles, and phylogenetic preparation. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.
Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.
How to run multiple sequence alignment online
Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.
- Curate sequences. Remove duplicates, obvious fragments, contaminants, and incompatible domain architectures.
- Choose methods. Select methods appropriate for sequence count, length, divergence, and long insertions.
- Run alignments. Run each method with recorded settings and preserve native output order.
- Compare columns. Review conserved blocks, gap-rich regions, outliers, and method disagreement.
- Export downstream set. Export the chosen alignment with exclusions, masks, and provenance.
How to interpret multiple sequence alignment results
Conserved columns are meaningful only relative to the sampled family and alignment quality. Review residue composition, sequence coverage, gap frequency, and whether apparent conservation is driven by many nearly identical records.
Before phylogenetics, remove nonhomologous flanks and consider masking unstable columns using a documented rule. Never hand-edit an alignment without retaining the original and recording each change.
How multiple sequence alignment works
Run Clustal Omega, MAFFT L-INS-i, and MUSCLE5 from one FASTA set and compare their native alignments.
- Curate sequences. Remove duplicates, obvious fragments, contaminants, and incompatible domain architectures.
- Choose methods. Select methods appropriate for sequence count, length, divergence, and long insertions.
- Run alignments. Run each method with recorded settings and preserve native output order.
- Compare columns. Review conserved blocks, gap-rich regions, outliers, and method disagreement.
- Export downstream set. Export the chosen alignment with exclusions, masks, and provenance.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Alignment input.
FASTAPDBmmCIFA curated homologous protein, DNA, or RNA FASTA set.
Outputs
- Alignment outputs.
FASTACSVTSVPDBJSONMethod-specific aligned FASTA or Clustal files, guide information, and review-ready exports.
Tools for multiple sequence alignment
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

Clustal Omega
Create scalable protein or nucleotide multiple-sequence alignments

MAFFT
Create configurable protein, DNA, or RNA alignments

MUSCLE5
Generate conventional or ensemble multiple-sequence alignments

HMMER
Find homologs with profile hidden Markov models

MMseqs2
Search and cluster large protein or nucleotide sequence sets

FastTree
Estimate trees from large sequence alignments

IQ-TREE
Infer maximum-likelihood phylogenies from alignments

RAxML-NG
Run maximum-likelihood phylogenetic analysis

RNAalifold
Predict consensus RNA structure from an RNA alignment

StringZilla v5
Calculate pairwise global, local, or edit-distance score matrices

FoldSeek
Find and compare structurally similar proteins

USAlign
Align macromolecular structures and derive residue correspondence
Other sequence analysis workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Pairwise sequence alignment
Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.
Global sequence alignment
Compares sequences end to end, including terminal differences and gaps across their full lengths.
Protein sequence alignment
Aligns amino-acid sequences using substitution-aware methods suited to protein evolution and function.
DNA sequence alignment
Aligns nucleotide sequences to compare homologous genes, amplicons, loci, transcripts, or constructs.
Local sequence alignment
Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.
Whole genome alignment
Maps large homologous regions between genome assemblies and reports coordinates, rearrangements, and sequence differences.
Structure-based sequence alignment
Uses three-dimensional correspondence to align residues whose sequence similarity alone may be weak.
Frequently asked questions
A curated homologous protein, DNA, or RNA FASTA set.
Method-specific aligned FASTA or Clustal files, guide information, and review-ready exports.
Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.
No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.
Retain the unaligned input, every method and setting, excluded sequences, masks, and the final alignment used downstream.
A complete multiple sequence alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.
The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured multiple sequence alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.