ProteinIQ
Sign inStart for free
ProteinIQ

Sequence analysis

Structure-based sequence alignment

Use three-dimensional geometry to derive residue correspondence when sequence similarity alone is too weak or ambiguous.

Open workflowCompare alignment types
Structure-based sequence alignmentWorkflow preview

Inputs

2 required

Methods

1 connected

  1. 01USAlign · Structure-Derived Residue Alignment

USAlign superposes mobile and reference structures and derives a residue alignment from their three-dimensional correspondence.

Use this template

On this page

  • Overview
  • Methods
  • Applications
  • Online workflow
  • Interpretation
  • How it works
  • Inputs & outputs

What is structure-based sequence alignment?

Structure-based sequence alignment is the process of arranging protein sequences by matching residues that occupy corresponding positions in three-dimensional structures. It is especially useful for remote homologs whose folds are conserved despite low sequence identity. In ProteinIQ, USAlign performs structural superposition and returns the resulting residue alignment, TM-scores, RMSD, sequence identity, and superposed coordinates.

Choose structure-based alignment when compatible experimental or predicted structures exist and sequence-only methods disagree or fail to recover a conserved fold. The method can clarify core correspondence across remote homologs, circular permutations, oligomers, and nucleic-acid structures, depending on the selected USAlign mode.

Structural alignment is not independent of structure quality. Missing residues, alternate conformations, domain movements, chain selection, prediction uncertainty, and different biological assemblies can all change the correspondence. Review coverage, both normalized TM-scores, RMSD, aligned length, sequence identity, and the superposed coordinates together.

When to use structure-based sequence alignment

  • Best fit. Remote homologs, conserved folds, and structure-guided residue mapping
  • Required input. Two compatible PDB or mmCIF structures with correct chain and assembly choices

Benefits of structure-based sequence alignment

  • Clear correspondence. Recovers correspondence at low identity
  • Connected evidence. Connects alignment to 3D geometry
  • Reusable output. Supports proteins and nucleic acids

Primary limitations

  • Method dependence. Depends on structure quality
  • Input dependence. Flexible domains can dominate RMSD
  • Interpretive limit. Structural similarity does not prove function

Structure-based sequence alignment methods

USAlign searches for a residue correspondence and rigid-body superposition that optimize a length-normalized TM-score. TM-score and RMSD summarize different aspects of the result and should not be treated as interchangeable.

FoldSeek can identify structurally similar candidates at scale, while USAlign is suited to detailed pairwise superposition. Sequence extraction may help compare structure-derived and sequence-only alignments without conflating them.

Structure-based sequence alignment applications

Structure-based sequence alignment is best suited to remote homologs, conserved folds, and structure-guided residue mapping. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run structure-based sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Choose structures. Choose structures representing the intended state, construct, and biological assembly.
  2. Select chains. Set chains, molecule type, oligomer handling, and any circular-permutation mode.
  3. Run USAlign. Run USAlign and preserve both score normalizations and aligned coordinates.
  4. Inspect superposition. Inspect coverage, RMSD, gaps, flexible regions, and the superposed model.
  5. Transfer cautiously. Transfer residue annotations only across well-supported structural correspondence.

How to interpret structure-based sequence alignment results

Inspect aligned length and both TM-score normalizations because unequal chain lengths can produce asymmetric values. RMSD should always be read with the number and fraction of aligned residues.

Do not transfer catalytic, binding, or numbering annotations across a gap or poorly superposed loop without local evidence. Conserved global folds can support very different biochemical roles.

How structure-based sequence alignment works

USAlign superposes mobile and reference structures and derives a residue alignment from their three-dimensional correspondence.

  1. Choose structures. Choose structures representing the intended state, construct, and biological assembly.
  2. Select chains. Set chains, molecule type, oligomer handling, and any circular-permutation mode.
  3. Run USAlign. Run USAlign and preserve both score normalizations and aligned coordinates.
  4. Inspect superposition. Inspect coverage, RMSD, gaps, flexible regions, and the superposed model.
  5. Transfer cautiously. Transfer residue annotations only across well-supported structural correspondence.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Two protein, RNA, or DNA structures in PDB or mmCIF format.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Structure-derived residue alignment, TM-scores, RMSD, aligned length, identity, and superposed coordinates.

Tools for structure-based sequence alignment

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

USAlign

Align macromolecular structures and derive residue correspondence

FoldSeek

Find and compare structurally similar proteins

PDB to FASTA converter

Extract sequences from structures before comparison

MAFFT

Create configurable protein, DNA, or RNA alignments

Clustal Omega

Create scalable protein or nucleotide multiple-sequence alignments

MUSCLE5

Generate conventional or ensemble multiple-sequence alignments

HMMER

Find homologs with profile hidden Markov models

MMseqs2

Search and cluster large protein or nucleotide sequence sets

StringZilla v5

Calculate pairwise global, local, or edit-distance score matrices

FastTree

Estimate trees from large sequence alignments

IQ-TREE

Infer maximum-likelihood phylogenies from alignments

RAxML-NG

Run maximum-likelihood phylogenetic analysis

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Pairwise sequence alignment

Compares two biological sequences and reports their residue-to-residue correspondence or a defined pairwise score.

Multiple sequence alignment

Aligns three or more homologous sequences to identify shared positions, insertions, deletions, and conserved regions.

Global sequence alignment

Compares sequences end to end, including terminal differences and gaps across their full lengths.

Protein sequence alignment

Aligns amino-acid sequences using substitution-aware methods suited to protein evolution and function.

DNA sequence alignment

Aligns nucleotide sequences to compare homologous genes, amplicons, loci, transcripts, or constructs.

Local sequence alignment

Finds or scores the best-matching subsequences without forcing unrelated flanks into the comparison.

Whole genome alignment

Maps large homologous regions between genome assemblies and reports coordinates, rearrangements, and sequence differences.

Frequently asked questions

Two protein, RNA, or DNA structures in PDB or mmCIF format.

Structure-derived residue alignment, TM-scores, RMSD, aligned length, identity, and superposed coordinates.

Start from the scientific scope: global or local, pairwise or multiple, sequence or structure, and conventional or genome scale. Then record the method, substitution model, gap settings, sequence type, and any filtering rather than relying on defaults without provenance.

No. Scores and identities quantify similarity under a defined model. Homology is an evolutionary interpretation, and shared function requires additional evidence such as domain context, conserved residues, structure, phylogeny, experiments, or curated annotation.

Retain structure provenance, chains, assembly, alignment mode, both TM-scores, RMSD, aligned length, and superposed coordinates.

A complete structure-based sequence alignment project is usually quote-based because providers scope sequence curation, method selection, alignment review, interpretation, and downstream analysis together. Harvard’s FY26 bioinformatics core first defines deliverables and a time estimate, then charges $180–$265 per hour; MSU lists $84–$110 per hour and expects at least eight consultant hours for custom analysis.

The total depends on sequence count and length, input cleanup, molecular type, the number of methods compared, manual review, genome scale, figures, phylogenetic or structural follow-up, and whether the deliverable includes interpretation or only alignment files.

ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured structure-based sequence alignment run estimated in credits before submission. Done-for-you analysis is scoped separately and can include data preparation, method comparison, interpretation, and a reproducible handoff.

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow
ProteinIQ

© 2026 ProteinIQ

Products

  • Bioinformatics tools
  • Workflows
  • PDB viewer
  • API

Solutions

  • Small molecule
  • RNA discovery
  • Antibody engineering
  • Peptide discovery
  • Enzyme engineering
  • Protein engineering
  • Virtual screening
  • Molecular docking
  • Protein structure prediction
  • RNA structure prediction
  • Protein structure alignment
  • Protein design
  • Sequence alignment
  • Phylogenetic analysis
  • Molecular dynamics simulation

Resources

  • Documentation
  • Blog
  • Guides
  • Datasets
  • Changelog
  • Sitemap

Company

  • About
  • Contact
  • Enterprise
  • Pricing
  • Security
  • Trust center
  • Author
  • Legal
  • Terms
  • Privacy policy

Connect

  • LinkedIn
  • X
  • Discord
  • Pricing