USAlign icon

USAlign

(20260527)Code (opens in a new tab)Paper (opens in a new tab)Docs

Universal structure alignment for proteins, RNA, and DNA molecules

Input

Upload file or drag and dropPDB, ENT, CIF, PDB.GZ, ENT.GZ, CIF.GZ, BZ2, TXT, XYZ · up to 50 MB
Upload file or drag and dropPDB, ENT, CIF, PDB.GZ, ENT.GZ, CIF.GZ, BZ2, TXT, XYZ · up to 50 MB

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job”.

Compare protein, RNA and DNA structures

USAlign compares three-dimensional macromolecular structures and reports native TM-scores, RMSD in angstrom, aligned residue counts and sequence identity. ProteinIQ runs USAlign version 20260527 and keeps its reported score normalizations and output files.

Pricing

Pricing starts at 11 credits. The calculator uses a 10-credit base plus 1 credit per detected PDB structure block. Multiple models can increase the count; inputs without PDB structure markers use a content-based fallback. The exact quote is calculated before submission.

Example inputCredits
Two PDB files, each containing one model ending with END12
Three PDB files in a structure collection, each containing one model ending with END13
Two mmCIF files without PDB structure markers11

Alignment mode, fast alignment, normalization and output settings do not add a separate multiplier to the current quote. The examples assume the file contents described above, rather than a fixed price for every comparison.

Inputs

For a pairwise comparison, provide Structure 1 (Mobile) and Structure 2 (Reference). The mobile structure is transformed onto the fixed reference. Upload PDB, ENT or mmCIF files, or retrieve structures using RCSB PDB codes. Each uploaded file can be up to 50 MiB.

Gzip-compressed PDB, ENT and mmCIF files and bzip2-compressed coordinate files are accepted without changing the captured bytes. SPICKER and USAlign XYZ files require an explicit format choice for the corresponding structure. These native coordinate-list formats can produce alignment scores but cannot use the native superposed-coordinate writer.

InputAccepted filesLimit or requirement
Structure 1 (Mobile), Structure 2 (Reference), Structure collection.pdb, .ent, .cif, .pdb.gz, .ent.gz, .cif.gz, .bz2, .txt, .xyz50 MiB per file; protein, RNA, DNA or mixed structures; structure inputs also support RCSB PDB retrieval
Initial or final FASTA alignment.fasta, .ali, .txt50 MiB; a two-sequence alignment for the supplied-alignment arrangement
Chain mapping file.txt, .tsv50 MiB; native two-column, tab-separated chain IDs for oligomer alignment
Structure pair list.txt, .tsv50 MiB; two uploaded structure filenames per row

Choose an Input arrangement for additional file roles:

ArrangementFiles
Compare two structuresMobile and reference structures
Use a supplied alignmentTwo structures and a two-sequence FASTA alignment; choose initial alignment (-i) or final alignment (-I)
Use a chain mappingTwo complexes and the native chain mapping text file
Align monomers to a complexA collection of monomer files and one reference complex
Multiple structure alignmentA collection of structures
All against allA collection of structures, compared by the native list operation
Query against a structure listOne mobile query and a structure collection
Structure list against a referenceA structure collection and one fixed reference
Explicit structure pairsA structure collection and a text file containing two uploaded filenames per row

For explicit pairs, use unique filenames containing letters, digits, dots, underscores, plus or minus signs. Pair-list entries must match uploaded filenames exactly. Collection order is preserved. Directory comparison modes report native results without a superposed coordinate file, except native multiple-structure and monomer-to-complex alignment. The worker limit is five minutes for the whole job, so large collections can time out.

Settings

Labels below match the form. Enum defaults show the saved value and visible option label. Empty selectors leave the native choice unchanged; switches default to off unless stated otherwise.

Input arrangement

ParameterTypeDefaultDescription
Inputsenumpair (Compare two structures)Selects one of the nine input arrangements listed above and its required file roles.
Supplied alignmentenuminitial (Initial alignment (-i))initial starts the native search from the uploaded alignment; fixed (Final alignment (-I)) uses it as the final alignment. Applies to Use a supplied alignment.

Alignment options

ParameterTypeDefaultDescription
Alignment modeenummonomer (Monomer (default))Selects the native algorithm from the mode table below.
Chain handlingenumauto (Native default for the selected mode)0: All chains, all models (biological unit); 1: All chains, first model (asymmetric unit); 2: First chain only (default); 3: First chain segment (up to TER record). The form default is auto, which uses all chains of the first model for complexes and applicable residue mappings, and the first chain otherwise.
Molecule typeenumauto (Auto-detect (default))prot: Protein only; RNA: RNA/DNA only. Auto-detection accepts both protein and nucleic acid residues, including mixed structures.
HETATM residuesenum0 (Ignore HETATM (default))1: Include HETATM residues; 2: Include MSE residues only, alongside ATOM residues. Mirror comparisons recommend including eligible HETATM residues for D amino acids, but preserve the selected value.
Fast alignmentbooleanfalseUses fTM-align for faster, potentially less accurate alignment.
Alignment mode optionValueNative behavior
Monomer (default)monomerAligns monomeric structures or individual chain pairs.
Oligomer (multi-chain)oligomerAligns multi-chain complexes.
Circular permutationcircularAligns structures with reordered termini.
Mirror imagemirrorMirrors the mobile structure before monomer alignment.
Monomers to complexmonomersAligns a collection of monomers to a reference complex; requires Align monomers to a complex inputs.
Multiple structure alignmentmultipleForms a consensus alignment from a collection; requires Multiple structure alignment inputs.
Fully non-sequentialnonsequentialUses the native fully non-sequential algorithm.
Semi-non-sequentialsemisequentialUses the native semi-non-sequential algorithm.
Flexible alignmentflexibleUses the native flexible-alignment algorithm with the selected hinge limit.

Chain selection

ParameterTypeDefaultDescription
Structure 1 chainsstringoptional (empty)Comma-separated mobile chain IDs, such as A,B; _ represents a blank chain ID. Empty uses all eligible chains subject to Chain handling.
Structure 2 chainsstringoptional (empty)Comma-separated reference chain IDs with the same syntax and chain-handling restriction.

Advanced native options

ParameterTypeDefaultDescription
Mirror the mobile structurebooleanfalseEnables native mirroring in other alignment modes; Mirror image alignment enables mirroring regardless of this switch.
Structure 1 modelsstringoptional (empty)Comma-separated native mobile MODEL IDs, such as 1,2; empty uses the native default. Chain handling still governs which models are eligible.
Structure 2 modelsstringoptional (empty)Comma-separated reference MODEL IDs with the same syntax and chain-handling restriction.
Representative atomstringautoRequired, 1 to 4 printable ASCII characters; spaces are preserved. auto selects protein CA or nucleic acid C3'; PC4' selects the native hybrid P/C4' rule.
Split structuresenumauto (Native default)0: Keep as one chain; 1: Split by MODEL; 2: Split by chain. The native default keeps one chain for residue-equivalence mode 2 and splits by chain otherwise.
Structure 1 formatenum-1 (Auto-detect PDB/mmCIF)0: PDB; 1: SPICKER; 2: USAlign XYZ; 3: mmCIF. SPICKER and XYZ require an explicit choice.
Structure 2 formatenum-1 (Auto-detect PDB/mmCIF)Uses the same format choices for the reference input.
Residue equivalenceenum0 (Structural alignment)Native residue correspondence modes 0 through 7 are listed below.
Sequence-dependent alignment (-seq)booleanfalseNative alias for Glocal sequence alignment, residue-equivalence mode 5; Residue equivalence must be 0 or 5.
Evaluate existing superposition (-se)booleanfalseExtracts an alignment without rotating or translating the input structures.
Nearest atoms for initial non-sequential alignmentinteger-1Minimum -1; that value preserves the native default of 5 for Fully non-sequential and 0 otherwise. A nonnegative value selects the native nearest-atom initialization count.
Maximum flexible hingesinteger9Range 0 to 9; applies to Flexible alignment.
Residue equivalence optionValue
Structural alignment0
Same residue index1
Same residue index and chain ID2
Same residue index and chain order3
Global sequence alignment4
Glocal sequence alignment5
Same residue index after chain mapping6
Sequence alignment after chain mapping7

Score normalization

ParameterTypeDefaultDescription
Additional TM-score normalizationenumnative (Native default)T: Average length; F: Second structure; -1: Shorter length; -2: Longer length; 1: Average length (numeric). Additional scores retain their native normalization labels.
Assigned normalization length (-u)number0Minimum 0; 0 omits the option. A positive length adds a native score without changing the alignment. An assigned length shorter than the structure can produce a score above 1.
Assigned d0 in angstrom (-d)number0Minimum 0; 0 omits the option. A positive distance adds the native distance-scaled score without changing the alignment.
TM-score early-stop cutoffnumber-1-1 disables early stopping. Native guidance is 0.5 to below 1; unlikely comparisons can stop early. The cutoff uses the selected additional normalization.

Native outputs

ParameterTypeDefaultDescription
Coordinate and visualization filesenumpymol (PyMOL)chimerax: ChimeraX; rasmol: RasMol; none: Report only. Coordinates are produced only for supported input arrangements and formats.
Save rotation matrixbooleanfalseSaves the native rotation matrix when supported; list-comparison modes print matrices in the native report.
Report aligned residue distancesbooleanfalseAdds the native aligned-residue distance table to the report.
Include individual alignments in complex or multiple resultsbooleanfalseRequests native full individual alignments, particularly for Monomers to complex and Multiple structure alignment; requires a multimer alignment mode.
Native report formatenum0 (Full report)1: Compact FASTA; 2: Compact tabular; -1: Full report without citation header. Compact tabular restricts additional normalization as described below.

Compatible combinations

Incompatible settings are rejected rather than silently changed:

  • Oligomer (multi-chain) and Monomers to complex require Chain handling auto, 0 or 1. Residue equivalence 2, 3 and 6 have the same chain-handling requirement.
  • Use a chain mapping requires Oligomer (multi-chain). Use a supplied alignment requires Monomer (default) or Mirror image, Residue equivalence 0, and Sequence-dependent alignment (-seq) off.
  • Align monomers to a complex and Multiple structure alignment inputs require their matching alignment modes, and those modes require the matching input arrangements.
  • Oligomer (multi-chain), Monomers to complex, Multiple structure alignment, Fully non-sequential, Semi-non-sequential and Flexible alignment require Assigned normalization length (-u) 0, Residue equivalence 0, and Sequence-dependent alignment (-seq) off.
  • Compact tabular requires Additional TM-score normalization native or F, Assigned normalization length (-u) 0, and Assigned d0 in angstrom (-d) 0.
  • Include individual alignments in complex or multiple results requires Oligomer (multi-chain), Monomers to complex, Multiple structure alignment, Fully non-sequential, Semi-non-sequential or Flexible alignment.

Job details

ParameterTypeDefaultDescription
Job namestringoptional (empty)Labels the job in its history; it does not change the scientific calculation.

Scores and files

The data table retains every native alignment record in source order, including structure/chain identifiers and every reported TM-score normalization. Sequence identity is a fraction, such as 0.25 for 25%. RMSD is in angstrom. A score normalized by the fixed reference length is useful when comparing a prediction with an experimental reference. Always compare scores using the same normalization.

Additional normalization options report scores using the average, shorter or longer structure length, an assigned length (-u), or an assigned distance scale in angstrom (-d). These options do not change the final native alignment. An assigned length shorter than the structure can produce a TM-score above 1. The native TM-score cutoff can stop unlikely comparisons early.

Choose PyMOL, ChimeraX, RasMol or report-only output. Downloads retain the native coordinate variants and original scripts, a complete native report, separate stdout, stderr when present, and JSON provenance with the command, source revision, captured-input hashes and native-file hashes. PyMOL and ChimeraX scripts also have additional portable copies that reference download filenames. Enable the rotation matrix or aligned-residue distance table when needed. Compact FASTA and tabular reports preserve their native text.

When monomer alignment processes several chain pairs, USAlign writes the superposition for the last pair. Coordinate metadata identifies that record, while the table retains all pairs. Multi-structure or complex reports can have several records without a single per-record coordinate association; the complete native report remains available.

Workflows provide the aligned structure, metric table and a collection of reports, scripts and supporting files, including rotation matrices when produced. A scoring-only native operation explicitly omits the aligned structure.

Source

See the USAlign source and native command-line documentation for file syntax and algorithm details. Historical jobs retain their admitted execution behavior; their stored settings and provenance identify that behavior.

Table of contents

Related tools

MAFFT

MAFFT

Align protein or nucleotide sequences with selectable accuracy and speed trade-offs.

MUSCLE5

MUSCLE5

Align multiple protein or nucleotide sequences with high-accuracy PPP refinement.

StringZilla v5

StringZilla v5

Hardware-accelerated edit distances and global or local sequence scores

Clustal Omega

Clustal Omega

Align multiple protein or nucleotide sequences and export FASTA, Clustal, or Phylip outputs.

FastTree

FastTree

Build phylogenetic trees from aligned protein or nucleotide sequences using approximate maximum-likelihood methods.

FoldSeek

FoldSeek

Search AlphaFold DB, compare structures, or cluster by 3D similarity

IQ-TREE

IQ-TREE

Build maximum likelihood phylogenetic trees with automatic model selection and ultrafast bootstrap.

MMseqs2

MMseqs2

Search and cluster protein or nucleotide sequences for homology discovery at large scale.

MUMmer4

MUMmer4

Align and compare whole genomes to detect SNPs, indels, and structural variants.

DockQ

DockQ

Assess docking quality between model and native structures