Align multiple protein or nucleotide sequences and export FASTA, Clustal, or Phylip outputs. Learn more
Run
0/1,000,000
Configuration
5 credits
Output
Configure inputs to begin
Set options on the left, then click “Align Sequences” — or start from an example.
Vertebrate beta-globins · refined protein MSA
E. coli lac operators · exact distances and guide tree
Human let-7 family · refined RNA alignment
Clustal Omega
(1.2.4)
Align multiple protein or nucleotide sequences and export FASTA, Clustal, or Phylip outputs. Learn more
Run
0/1,000,000
Configuration
5 credits
Output
Configure inputs to begin
Set options on the left, then click “Align Sequences” — or start from an example.
Vertebrate beta-globins · refined protein MSA
E. coli lac operators · exact distances and guide tree
Human let-7 family · refined RNA alignment
What is Clustal Omega?
Clustal Omega is a multiple sequence alignment (MSA) program that aligns protein or nucleotide sequences to reveal conserved regions, evolutionary relationships, and functional motifs. MSA is a foundational step in many bioinformatics workflows—from phylogenetic analysis to structure prediction to primer design.
Clustal Omega can handle datasets ranging from a handful of sequences to tens of thousands, making it suitable for both focused studies and large-scale comparative genomics. Once you have an alignment, you can use FastTree to build a phylogenetic tree from it.
How does Clustal Omega work?
Clustal Omega uses a progressive alignment strategy: it first estimates how similar sequences are to each other, builds a guide tree from those similarities, then aligns sequences following the tree order. The key innovations are the mBed algorithm for scalability and HMM-based profile alignment for accuracy.
The mBed algorithm
Traditional pairwise distance calculation scales as O(N2), which becomes prohibitive for large datasets. The mBed algorithm reduces this to O(NlogN) by "embedding" each sequence into a low-dimensional space.
Instead of comparing every sequence to every other sequence, mBed selects a small set of reference sequences and represents each sequence as a vector of distances to these references. These vectors can be clustered rapidly using k-means, with clusters capped at 100 sequences. Full distance matrices are only computed within clusters, not across the entire dataset.
HMM-based profile alignment
When combining two groups of aligned sequences (profiles), Clustal Omega uses hidden Markov model alignment via the HHalign package. Each profile is converted to an HMM with match, insert, and delete states. Aligning two HMMs rather than simple position-specific scoring matrices improves sensitivity for distantly related sequences.
Guide tree and progressive alignment
The distance matrix (partial or full) is used to construct a guide tree via UPGMA. This tree determines the order of pairwise alignments: closely related sequences are aligned first, then progressively merged with more distant groups until all sequences are incorporated.
Alignment settings
Sequence type
Clustal Omega auto-detects whether your sequences are protein or nucleotide by examining character composition. Manual selection (Protein, DNA, or RNA) is useful when auto-detection might be ambiguous—for example, with very short sequences or sequences containing unusual characters.
Output format
FASTA: Aligned sequences with gaps represented as -. Most compatible with downstream tools including FastTree.
Clustal: Traditional format showing alignment blocks with conservation symbols. Good for visual inspection.
Phylip: Fixed-width format used by phylogenetic programs.
MSF: GCG format, useful for legacy software.
Stockholm: Annotated format used by Pfam and Rfam databases.
Refinement iterations
After the initial alignment, Clustal Omega can refine it by rebuilding the guide tree from the alignment itself (rather than pairwise distances) and realigning. Each iteration uses the preceding alignment to construct the next guide tree.
ProteinIQ supports integer values from 0 to 5. Use 1-2 iterations when you want to evaluate whether refinement helps your dataset. For exploratory work or very large datasets, skip refinement (0 iterations, the native default) to save time.
Full distance matrix
By default, mBed approximates the information needed for scalable guide-tree construction. Enabling the full distance matrix computes every pairwise distance at the cost of O(N2) time and memory.
The full matrix is useful when you specifically need all pairwise distances or want to compare guide-tree construction methods. It does not guarantee a better biological alignment. For large datasets, mBed is usually the practical choice.
Output order
Controls whether sequences in the output follow your original input order or are reordered by the guide tree. Tree order groups similar sequences together, which can make visual inspection easier and is the natural order for phylogenetic workflows. Input order (the default) preserves whatever order you provided.
Remove existing alignment
Clustal Omega can use an existing input alignment to build an HMM before it realigns the underlying sequences. Enable this option when you want to discard that existing alignment information and realign the ungapped sequences from scratch.
Output distance matrix
Produces an additional output file containing Clustal Omega's native uncorrected pair distances. Values range from 0 for identical aligned residues to 1 when no aligned residues are identical. These are operational sequence distances, not Kimura-corrected evolutionary distances.
Output guide tree
Produces an additional Newick-format tree file showing the guide tree used to order the progressive alignment. This is the UPGMA tree built from pairwise distances (or mBed-approximated distances). It is an operational alignment tree, not a phylogenetic inference.
Understanding the results
The output is a multiple sequence alignment where:
Columns represent homologous positions across sequences
Gap characters (-) indicate insertions or deletions relative to other sequences
Conserved columns (same residue across all sequences) suggest functional or structural importance
The alignment length reported is the number of columns, which will be longer than any individual sequence due to gaps. High-quality alignments have fewer scattered gaps and more continuous aligned blocks.
This example aligns beta-globin sequences from human, mouse, and chicken to show how Clustal Omega presents conserved positions and lineage-specific substitutions across homologous proteins.
Non-default settings:Sequence type = Protein, Output format = Clustal, and Refinement iterations = 1
Clustal Omega protein alignment of three vertebrate beta-globins with the conservation track and 147 alignment columns
The MSA viewer reports three sequences across 147 columns. Long conserved blocks are visible across the globin sequences, with substitutions distributed through the alignment. These shared columns are useful candidates for further functional or structural investigation, but the alignment alone does not establish residue function or a phylogenetic history.
This short-DNA example compares the submitted 5′→3′ lacO1, lacO2, and lacO3 operator strings. Manual DNA selection avoids ambiguity for 21-nucleotide inputs, while the additional outputs demonstrate the distance and guide-tree workflow.
Inputs: E. coli lacO1 primary operator and lacO2/lacO3 auxiliary operators
Non-default settings:Sequence type = DNA, Use full distance matrix = enabled, Output order = Tree order, Output distance matrix = enabled, and Output guide tree = enabled
Clustal Omega DNA alignment of the lacO1, lacO2, and lacO3 operators in tree order across 21 columns
The 21-column alignment groups the three operators in tree order and makes their conserved core pattern and substitutions directly inspectable. The output order reflects Clustal Omega's operational guide tree; it is not evidence of regulatory strength or evolutionary ancestry.
Clustal Omega files panel listing the lac operator FASTA alignment, pair-distance matrix, and Newick guide tree
The Files view confirms that the run returned the aligned FASTA, the native uncorrected pair-distance matrix, and the Newick guide tree. The matrix values measure sequence difference in this alignment, not binding affinity, and the UPGMA tree is an alignment aid rather than a phylogenetic inference.
This example aligns five mature human let-7 5p sequences to demonstrate explicit RNA handling, short-sequence conservation, and Stockholm output for downstream RNA-family workflows.
Inputs: miRBase MIMAT0000062, MIMAT0000063, MIMAT0000064, MIMAT0000065, and MIMAT0000066
Non-default settings:Sequence type = RNA, Output format = Stockholm, and Refinement iterations = 2
Clustal Omega RNA alignment of five human let-7 family members with the conservation track across 22 columns
All five mature miRNAs fit within a 22-column alignment, with a strongly conserved 5′ region and family-specific substitutions visible in the viewer. This pattern supports positional comparison within the submitted family, but it does not by itself prove shared targets, expression, or biological function.
Common workflows
Clustal Omega is typically the first step in a multi-tool pipeline:
Phylogenetic analysis: Align sequences with Clustal Omega → Build tree with FastTree
Conservation analysis: Align sequences → Identify conserved regions for mutagenesis targets
Homology modeling: Align target to templates → Use alignment for structure prediction
Primer design: Align variants → Design primers in conserved regions
Limitations
Clustal Omega assumes sequences are homologous and alignable. It will produce an alignment even for unrelated sequences, but the result will be meaningless. Always verify that your sequences share evolutionary or functional relationships before aligning.
Very divergent sequences (below ~20% identity for proteins) may not align reliably with any progressive alignment method. Consider structure-based alignment for such cases.
Clustal Omega is a multiple sequence alignment (MSA) program that aligns protein or nucleotide sequences to reveal conserved regions, evolutionary relationships, and functional motifs. MSA is a foundational step in many bioinformatics workflows—from phylogenetic analysis to structure prediction to primer design.
Clustal Omega can handle datasets ranging from a handful of sequences to tens of thousands, making it suitable for both focused studies and large-scale comparative genomics. Once you have an alignment, you can use FastTree to build a phylogenetic tree from it.
How does Clustal Omega work?
Clustal Omega uses a progressive alignment strategy: it first estimates how similar sequences are to each other, builds a guide tree from those similarities, then aligns sequences following the tree order. The key innovations are the mBed algorithm for scalability and HMM-based profile alignment for accuracy.
The mBed algorithm
Traditional pairwise distance calculation scales as O(N2), which becomes prohibitive for large datasets. The mBed algorithm reduces this to O(NlogN) by "embedding" each sequence into a low-dimensional space.
Instead of comparing every sequence to every other sequence, mBed selects a small set of reference sequences and represents each sequence as a vector of distances to these references. These vectors can be clustered rapidly using k-means, with clusters capped at 100 sequences. Full distance matrices are only computed within clusters, not across the entire dataset.
HMM-based profile alignment
When combining two groups of aligned sequences (profiles), Clustal Omega uses hidden Markov model alignment via the HHalign package. Each profile is converted to an HMM with match, insert, and delete states. Aligning two HMMs rather than simple position-specific scoring matrices improves sensitivity for distantly related sequences.
Guide tree and progressive alignment
The distance matrix (partial or full) is used to construct a guide tree via UPGMA. This tree determines the order of pairwise alignments: closely related sequences are aligned first, then progressively merged with more distant groups until all sequences are incorporated.
Alignment settings
Sequence type
Clustal Omega auto-detects whether your sequences are protein or nucleotide by examining character composition. Manual selection (Protein, DNA, or RNA) is useful when auto-detection might be ambiguous—for example, with very short sequences or sequences containing unusual characters.
Output format
FASTA: Aligned sequences with gaps represented as -. Most compatible with downstream tools including FastTree.
Clustal: Traditional format showing alignment blocks with conservation symbols. Good for visual inspection.
Phylip: Fixed-width format used by phylogenetic programs.
MSF: GCG format, useful for legacy software.
Stockholm: Annotated format used by Pfam and Rfam databases.
Refinement iterations
After the initial alignment, Clustal Omega can refine it by rebuilding the guide tree from the alignment itself (rather than pairwise distances) and realigning. Each iteration uses the preceding alignment to construct the next guide tree.
ProteinIQ supports integer values from 0 to 5. Use 1-2 iterations when you want to evaluate whether refinement helps your dataset. For exploratory work or very large datasets, skip refinement (0 iterations, the native default) to save time.
Full distance matrix
By default, mBed approximates the information needed for scalable guide-tree construction. Enabling the full distance matrix computes every pairwise distance at the cost of O(N2) time and memory.
The full matrix is useful when you specifically need all pairwise distances or want to compare guide-tree construction methods. It does not guarantee a better biological alignment. For large datasets, mBed is usually the practical choice.
Output order
Controls whether sequences in the output follow your original input order or are reordered by the guide tree. Tree order groups similar sequences together, which can make visual inspection easier and is the natural order for phylogenetic workflows. Input order (the default) preserves whatever order you provided.
Remove existing alignment
Clustal Omega can use an existing input alignment to build an HMM before it realigns the underlying sequences. Enable this option when you want to discard that existing alignment information and realign the ungapped sequences from scratch.
Output distance matrix
Produces an additional output file containing Clustal Omega's native uncorrected pair distances. Values range from 0 for identical aligned residues to 1 when no aligned residues are identical. These are operational sequence distances, not Kimura-corrected evolutionary distances.
Output guide tree
Produces an additional Newick-format tree file showing the guide tree used to order the progressive alignment. This is the UPGMA tree built from pairwise distances (or mBed-approximated distances). It is an operational alignment tree, not a phylogenetic inference.
Understanding the results
The output is a multiple sequence alignment where:
Columns represent homologous positions across sequences
Gap characters (-) indicate insertions or deletions relative to other sequences
Conserved columns (same residue across all sequences) suggest functional or structural importance
The alignment length reported is the number of columns, which will be longer than any individual sequence due to gaps. High-quality alignments have fewer scattered gaps and more continuous aligned blocks.
This example aligns beta-globin sequences from human, mouse, and chicken to show how Clustal Omega presents conserved positions and lineage-specific substitutions across homologous proteins.
Non-default settings:Sequence type = Protein, Output format = Clustal, and Refinement iterations = 1
Clustal Omega protein alignment of three vertebrate beta-globins with the conservation track and 147 alignment columns
The MSA viewer reports three sequences across 147 columns. Long conserved blocks are visible across the globin sequences, with substitutions distributed through the alignment. These shared columns are useful candidates for further functional or structural investigation, but the alignment alone does not establish residue function or a phylogenetic history.
This short-DNA example compares the submitted 5′→3′ lacO1, lacO2, and lacO3 operator strings. Manual DNA selection avoids ambiguity for 21-nucleotide inputs, while the additional outputs demonstrate the distance and guide-tree workflow.
Inputs: E. coli lacO1 primary operator and lacO2/lacO3 auxiliary operators
Non-default settings:Sequence type = DNA, Use full distance matrix = enabled, Output order = Tree order, Output distance matrix = enabled, and Output guide tree = enabled
Clustal Omega DNA alignment of the lacO1, lacO2, and lacO3 operators in tree order across 21 columns
The 21-column alignment groups the three operators in tree order and makes their conserved core pattern and substitutions directly inspectable. The output order reflects Clustal Omega's operational guide tree; it is not evidence of regulatory strength or evolutionary ancestry.
Clustal Omega files panel listing the lac operator FASTA alignment, pair-distance matrix, and Newick guide tree
The Files view confirms that the run returned the aligned FASTA, the native uncorrected pair-distance matrix, and the Newick guide tree. The matrix values measure sequence difference in this alignment, not binding affinity, and the UPGMA tree is an alignment aid rather than a phylogenetic inference.
This example aligns five mature human let-7 5p sequences to demonstrate explicit RNA handling, short-sequence conservation, and Stockholm output for downstream RNA-family workflows.
Inputs: miRBase MIMAT0000062, MIMAT0000063, MIMAT0000064, MIMAT0000065, and MIMAT0000066
Non-default settings:Sequence type = RNA, Output format = Stockholm, and Refinement iterations = 2
Clustal Omega RNA alignment of five human let-7 family members with the conservation track across 22 columns
All five mature miRNAs fit within a 22-column alignment, with a strongly conserved 5′ region and family-specific substitutions visible in the viewer. This pattern supports positional comparison within the submitted family, but it does not by itself prove shared targets, expression, or biological function.
Common workflows
Clustal Omega is typically the first step in a multi-tool pipeline:
Phylogenetic analysis: Align sequences with Clustal Omega → Build tree with FastTree
Conservation analysis: Align sequences → Identify conserved regions for mutagenesis targets
Homology modeling: Align target to templates → Use alignment for structure prediction
Primer design: Align variants → Design primers in conserved regions
Limitations
Clustal Omega assumes sequences are homologous and alignable. It will produce an alignment even for unrelated sequences, but the result will be meaningless. Always verify that your sequences share evolutionary or functional relationships before aligning.
Very divergent sequences (below ~20% identity for proteins) may not align reliably with any progressive alignment method. Consider structure-based alignment for such cases.