ProteinIQ
Sign inStart for free
ProteinIQ

Sequence analysis

Protein sequence alignment

Align amino-acid sequences with protein-aware methods and interpret conservation in structural and functional context.

Open workflowCompare alignment types
Protein sequence alignmentWorkflow preview

Inputs

1 required

Methods

3 connected

  1. 01MAFFT · Protein Alignment
  2. 02Clustal Omega
  3. 03MUSCLE5

Run protein-aware MAFFT, Clustal Omega, and MUSCLE5 alignments from the same amino-acid FASTA.

Use this template

On this page

  • Overview
  • Methods
  • Applications
  • Online workflow
  • Interpretation
  • How it works
  • Inputs & outputs

What is protein sequence alignment?

Protein sequence alignment is the process of arranging amino-acid sequences in rows so homologous or functionally comparable residues appear in shared columns. Protein methods use substitution models that reflect unequal evolutionary exchangeability among amino acids. Alignments can reveal conserved motifs, family-specific positions, insertions, deletions, and domain boundaries, but sequence correspondence remains a model-based hypothesis.

Start with proteins that plausibly share ancestry or architecture. Full-length proteins containing different domain combinations should often be split into comparable domains before alignment. Signal peptides, disordered tails, low-complexity regions, and repeat expansions can otherwise dominate gap placement and obscure conserved cores.

Method choice depends on sequence count, length, divergence, and expected insertions. MAFFT, Clustal Omega, and MUSCLE5 offer complementary strategies. For remote homologs, use HMMER, structural searches, or structure alignment to establish family membership and inspect difficult columns rather than treating one sequence-only alignment as definitive.

When to use protein sequence alignment

  • Best fit. Protein families, domains, motifs, variants, and engineered constructs
  • Required input. Protein FASTA sequences with validated translation and boundaries

Benefits of protein sequence alignment

  • Clear correspondence. Uses amino-acid substitution information
  • Connected evidence. Supports motif and family analysis
  • Reusable output. Connects sequence with structure and function

Primary limitations

  • Method dependence. Remote homology can be ambiguous
  • Input dependence. Domain mixtures distort columns
  • Interpretive limit. Conservation does not establish mechanism

Protein sequence alignment methods

Protein substitution matrices reward conservative exchanges differently from radical changes. Gap penalties model insertion and deletion events but cannot capture every evolutionary history.

Profiles and iterative refinement can improve family alignments. Structural correspondence is valuable when sequence identity is low, although structure-derived columns should be labeled separately from sequence-only results.

Protein sequence alignment applications

Protein sequence alignment is best suited to protein families, domains, motifs, variants, and engineered constructs. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run protein sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Curate proteins. Remove translation errors, duplicates, fragments, and unintended isoforms.
  2. Check architecture. Compare domain architecture and choose full-length or domain-level scope.
  3. Run protein methods. Run several protein-alignment methods with saved settings.
  4. Map annotations. Inspect motifs, gaps, conserved residues, outliers, and structural context.
  5. Export evidence. Export the chosen alignment with sequence and method provenance.

How to interpret protein sequence alignment results

Evaluate conserved chemistry, not only exact identity. A hydrophobic core position may tolerate several residues while a catalytic residue may require one side-chain chemistry.

Map annotations only across well-supported columns and preserve source evidence. Automated transfer across uncertain or gap-rich regions can propagate incorrect residue numbering and function.

How protein sequence alignment works

Run protein-aware MAFFT, Clustal Omega, and MUSCLE5 alignments from the same amino-acid FASTA.

  1. Curate proteins. Remove translation errors, duplicates, fragments, and unintended isoforms.
  2. Check architecture. Compare domain architecture and choose full-length or domain-level scope.
  3. Run protein methods. Run several protein-alignment methods with saved settings.
  4. Map annotations. Inspect motifs, gaps, conserved residues, outliers, and structural context.
  5. Export evidence. Export the chosen alignment with sequence and method provenance.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Translated protein FASTA records with correct boundaries and identifiers.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Aligned protein FASTA or Clustal files, method settings, and downstream-ready exports.

Tools for protein sequence alignment

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

MAFFT

Create configurable protein, DNA, or RNA alignments

Clustal Omega

Create scalable protein or nucleotide multiple-sequence alignments

MUSCLE5

Generate conventional or ensemble multiple-sequence alignments

HMMER

Find homologs with profile hidden Markov models

MMseqs2

Search and cluster large protein or nucleotide sequence sets

FoldSeek

Find and compare structurally similar proteins

USAlign

Align macromolecular structures and derive residue correspondence

PDB to FASTA converter

Extract sequences from structures before comparison

StringZilla v5

Calculate pairwise global, local, or edit-distance score matrices

FastTree

Estimate trees from large sequence alignments

IQ-TREE

Infer maximum-likelihood phylogenies from alignments

RAxML-NG

Run maximum-likelihood phylogenetic analysis

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Pairwise sequence alignment

Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.

Multiple sequence alignment

Aligns three or more homologous sequences to identify shared positions, insertions, deletions, and conserved regions.

Global sequence alignment

Compares sequences end to end, including terminal differences and gaps across their full lengths.

DNA sequence alignment

Aligns nucleotide sequences to compare homologous genes, amplicons, loci, transcripts, or constructs.

Local sequence alignment

Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.

Whole genome alignment

Maps large homologous regions between genome assemblies and reports coordinates, rearrangements, and sequence differences.

Structure-based sequence alignment

Uses three-dimensional correspondence to align residues whose sequence similarity alone may be weak.

Frequently asked questions

Translated protein FASTA records with correct boundaries and identifiers.

Aligned protein FASTA or Clustal files, method settings, and downstream-ready exports.

Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.

No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.

Retain translations, domain decisions, unaligned sequences, method settings, annotations, and any manually reviewed columns.

A complete protein sequence alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.

The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.

ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured protein sequence alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow
ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing