Use case

Structure-based sequence alignment

Use three-dimensional geometry to derive residue correspondence when sequence similarity alone is too weak or ambiguous.

Structure-based sequence alignmentRead-only preview

Inputs

2 required

Methods

1 connected

  1. 01USAlign · Structure-Derived Residue Alignment

USAlign superposes mobile and reference structures and derives a residue alignment from their three-dimensional correspondence.

Use this template

What is structure-based sequence alignment?

Structure-based sequence alignment is the process of arranging protein sequences by matching residues that occupy corresponding positions in three-dimensional structures. It is especially useful for remote homologs whose folds are conserved despite low sequence identity. In ProteinIQ, USAlign performs structural superposition and returns the resulting residue alignment, TM-scores, RMSD, sequence identity, and superposed coordinates.

Choose structure-based alignment when compatible experimental or predicted structures exist and sequence-only methods disagree or fail to recover a conserved fold. The method can clarify core correspondence across remote homologs, circular permutations, oligomers, and nucleic-acid structures, depending on the selected USAlign mode.

Structural alignment is not independent of structure quality. Missing residues, alternate conformations, domain movements, chain selection, prediction uncertainty, and different biological assemblies can all change the correspondence. Review coverage, both normalized TM-scores, RMSD, aligned length, sequence identity, and the superposed coordinates together.

When to use structure-based sequence alignment

  • Best fit. Remote homologs, conserved folds, and structure-guided residue mapping
  • Required input. Two compatible PDB or mmCIF structures with correct chain and assembly choices

Benefits of structure-based sequence alignment

  • Clear correspondence. Recovers correspondence at low identity
  • Connected evidence. Connects alignment to 3D geometry
  • Reusable output. Supports proteins and nucleic acids

Primary limitations

  • Method dependence. Depends on structure quality
  • Input dependence. Flexible domains can dominate RMSD
  • Interpretive limit. Structural similarity does not prove function

Structure-based sequence alignment methods

USAlign searches for a residue correspondence and rigid-body superposition that optimize a length-normalized TM-score. TM-score and RMSD summarize different aspects of the result and should not be treated as interchangeable.

FoldSeek can identify structurally similar candidates at scale, while USAlign is suited to detailed pairwise superposition. Sequence extraction may help compare structure-derived and sequence-only alignments without conflating them.

Structure-based sequence alignment applications

Structure-based sequence alignment is best suited to remote homologs, conserved folds, and structure-guided residue mapping. The result can support comparative review, sequence curation, annotation, profile construction, phylogenetic preparation, structural interpretation, or experimental planning when those downstream uses match the alignment scope.

Keep the alignment as evidence rather than a conclusion. Downstream claims should remain tied to sequence provenance, coverage, method agreement, relevant biological context, and any independent structural, evolutionary, or experimental support.

How to run structure-based sequence alignment online

Use the connected workflow to keep input records, method settings, native outputs, warnings, and exports together. Review every stage before using the result for annotation, phylogeny, variant interpretation, or experimental decisions.

  1. Choose structures. Choose structures representing the intended state, construct, and biological assembly.
  2. Select chains. Set chains, molecule type, oligomer handling, and any circular-permutation mode.
  3. Run USAlign. Run USAlign and preserve both score normalizations and aligned coordinates.
  4. Inspect superposition. Inspect coverage, RMSD, gaps, flexible regions, and the superposed model.
  5. Transfer cautiously. Transfer residue annotations only across well-supported structural correspondence.

How to interpret structure-based sequence alignment results

Inspect aligned length and both TM-score normalizations because unequal chain lengths can produce asymmetric values. RMSD should always be read with the number and fraction of aligned residues.

Do not transfer catalytic, binding, or numbering annotations across a gap or poorly superposed loop without local evidence. Conserved global folds can support very different biochemical roles.

How structure-based sequence alignment works

USAlign superposes mobile and reference structures and derives a residue alignment from their three-dimensional correspondence.

  1. Choose structures. Choose structures representing the intended state, construct, and biological assembly.
  2. Select chains. Set chains, molecule type, oligomer handling, and any circular-permutation mode.
  3. Run USAlign. Run USAlign and preserve both score normalizations and aligned coordinates.
  4. Inspect superposition. Inspect coverage, RMSD, gaps, flexible regions, and the superposed model.
  5. Transfer cautiously. Transfer residue annotations only across well-supported structural correspondence.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Alignment input. FASTA PDB mmCIF Two protein, RNA, or DNA structures in PDB or mmCIF format.

Outputs

  • Alignment outputs. FASTA CSV TSV PDB JSON Structure-derived residue alignment, TM-scores, RMSD, aligned length, identity, and superposed coordinates.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow