Use case

Protein sequence alignment

Align amino-acid sequences with protein-aware methods and interpret conservation in structural and functional context.

Protein sequence alignmentRead-only preview

Inputs

1 required

Methods

3 connected

  1. 01MAFFT · Protein Alignment
  2. 02Clustal Omega
  3. 03MUSCLE5

Run protein-aware MAFFT, Clustal Omega, and MUSCLE5 alignments from the same amino-acid FASTA.

Use this template

What is protein sequence alignment?

Protein sequence alignment is the process of arranging amino-acid sequences in rows so homologous or functionally comparable residues appear in shared columns. Protein methods use substitution models that reflect unequal evolutionary exchangeability among amino acids. Alignments can reveal conserved motifs, family-specific positions, insertions, deletions, and domain boundaries, but sequence correspondence remains a model-based hypothesis.

Start with proteins that plausibly share ancestry or architecture. Full-length proteins containing different domain combinations should often be split into comparable domains before alignment. Signal peptides, disordered tails, low-complexity regions, and repeat expansions can otherwise dominate gap placement and obscure conserved cores.

Method choice depends on sequence count, length, divergence, and expected insertions. MAFFT, Clustal Omega, and MUSCLE5 offer complementary strategies. For remote homologs, use HMMER, structural searches, or structure alignment to establish family membership and inspect difficult columns rather than treating one sequence-only alignment as definitive.

When to use protein sequence alignment

  • Best fit. Protein families, domains, motifs, variants, and engineered constructs
  • Required input. Protein FASTA sequences with validated translation and boundaries

Benefits of protein sequence alignment

  • Clear correspondence. Uses amino-acid substitution information
  • Connected evidence. Supports motif and family analysis
  • Reusable output. Connects sequence with structure and function

Primary limitations

  • Method dependence. Remote homology can be ambiguous
  • Input dependence. Domain mixtures distort columns
  • Interpretive limit. Conservation does not establish mechanism

Protein sequence alignment methods

Protein substitution matrices reward conservative exchanges differently from radical changes. Gap penalties model insertion and deletion events but cannot capture every evolutionary history.

Profiles and iterative refinement can improve family alignments. Structural correspondence is valuable when sequence identity is low, although structure-derived columns should be labeled separately from sequence-only results.

Protein sequence alignment applications

Protein sequence alignment is best suited to protein families, domains, motifs, variants, and engineered constructs. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run protein sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Curate proteins. Remove translation errors, duplicates, fragments, and unintended isoforms.
  2. Check architecture. Compare domain architecture and choose full-length or domain-level scope.
  3. Run protein methods. Run several protein-alignment methods with saved settings.
  4. Map annotations. Inspect motifs, gaps, conserved residues, outliers, and structural context.
  5. Export evidence. Export the chosen alignment with sequence and method provenance.

How to interpret protein sequence alignment results

Evaluate conserved chemistry, not only exact identity. A hydrophobic core position may tolerate several residues while a catalytic residue may require one side-chain chemistry.

Map annotations only across well-supported columns and preserve source evidence. Automated transfer across uncertain or gap-rich regions can propagate incorrect residue numbering and function.

How protein sequence alignment works

Run protein-aware MAFFT, Clustal Omega, and MUSCLE5 alignments from the same amino-acid FASTA.

  1. Curate proteins. Remove translation errors, duplicates, fragments, and unintended isoforms.
  2. Check architecture. Compare domain architecture and choose full-length or domain-level scope.
  3. Run protein methods. Run several protein-alignment methods with saved settings.
  4. Map annotations. Inspect motifs, gaps, conserved residues, outliers, and structural context.
  5. Export evidence. Export the chosen alignment with sequence and method provenance.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Translated protein FASTA records with correct boundaries and identifiers.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Aligned protein FASTA or Clustal files, method settings, and downstream-ready exports.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow