
USAlign
Universal structure alignment for proteins, RNA, and DNA molecules
Input
Compare protein, RNA and DNA structures
USAlign compares three-dimensional macromolecular structures and reports native TM-scores, RMSD in angstrom, aligned residue counts and sequence identity. ProteinIQ runs USAlign version 20260527 and keeps its reported score normalizations and output files.
Pricing
Pricing starts at 11 credits. The calculator uses a 10-credit base plus 1 credit per detected PDB structure block. Multiple models can increase the count; inputs without PDB structure markers use a content-based fallback. The exact quote is calculated before submission.
| Example input | Credits |
|---|---|
Two PDB files, each containing one model ending with END | 12 |
Three PDB files in a structure collection, each containing one model ending with END | 13 |
| Two mmCIF files without PDB structure markers | 11 |
Alignment mode, fast alignment, normalization and output settings do not add a separate multiplier to the current quote. The examples assume the file contents described above, rather than a fixed price for every comparison.
Inputs
For a pairwise comparison, provide Structure 1 (Mobile) and Structure 2 (Reference). The mobile structure is transformed onto the fixed reference. Upload PDB, ENT or mmCIF files, or retrieve structures using RCSB PDB codes. Each uploaded file can be up to 50 MiB.
Gzip-compressed PDB, ENT and mmCIF files and bzip2-compressed coordinate files are accepted without changing the captured bytes. SPICKER and USAlign XYZ files require an explicit format choice for the corresponding structure. These native coordinate-list formats can produce alignment scores but cannot use the native superposed-coordinate writer.
| Input | Accepted files | Limit or requirement |
|---|---|---|
| Structure 1 (Mobile), Structure 2 (Reference), Structure collection | .pdb, .ent, .cif, .pdb.gz, .ent.gz, .cif.gz, .bz2, .txt, .xyz | 50 MiB per file; protein, RNA, DNA or mixed structures; structure inputs also support RCSB PDB retrieval |
| Initial or final FASTA alignment | .fasta, .ali, .txt | 50 MiB; a two-sequence alignment for the supplied-alignment arrangement |
| Chain mapping file | .txt, .tsv | 50 MiB; native two-column, tab-separated chain IDs for oligomer alignment |
| Structure pair list | .txt, .tsv | 50 MiB; two uploaded structure filenames per row |
Choose an Input arrangement for additional file roles:
| Arrangement | Files |
|---|---|
| Compare two structures | Mobile and reference structures |
| Use a supplied alignment | Two structures and a two-sequence FASTA alignment; choose initial alignment (-i) or final alignment (-I) |
| Use a chain mapping | Two complexes and the native chain mapping text file |
| Align monomers to a complex | A collection of monomer files and one reference complex |
| Multiple structure alignment | A collection of structures |
| All against all | A collection of structures, compared by the native list operation |
| Query against a structure list | One mobile query and a structure collection |
| Structure list against a reference | A structure collection and one fixed reference |
| Explicit structure pairs | A structure collection and a text file containing two uploaded filenames per row |
For explicit pairs, use unique filenames containing letters, digits, dots, underscores, plus or minus signs. Pair-list entries must match uploaded filenames exactly. Collection order is preserved. Directory comparison modes report native results without a superposed coordinate file, except native multiple-structure and monomer-to-complex alignment. The worker limit is five minutes for the whole job, so large collections can time out.
Settings
Labels below match the form. Enum defaults show the saved value and visible option label. Empty selectors leave the native choice unchanged; switches default to off unless stated otherwise.
Input arrangement
| Parameter | Type | Default | Description |
|---|---|---|---|
| Inputs | enum | pair (Compare two structures) | Selects one of the nine input arrangements listed above and its required file roles. |
| Supplied alignment | enum | initial (Initial alignment (-i)) | initial starts the native search from the uploaded alignment; fixed (Final alignment (-I)) uses it as the final alignment. Applies to Use a supplied alignment. |
Alignment options
| Parameter | Type | Default | Description |
|---|---|---|---|
| Alignment mode | enum | monomer (Monomer (default)) | Selects the native algorithm from the mode table below. |
| Chain handling | enum | auto (Native default for the selected mode) | 0: All chains, all models (biological unit); 1: All chains, first model (asymmetric unit); 2: First chain only (default); 3: First chain segment (up to TER record). The form default is auto, which uses all chains of the first model for complexes and applicable residue mappings, and the first chain otherwise. |
| Molecule type | enum | auto (Auto-detect (default)) | prot: Protein only; RNA: RNA/DNA only. Auto-detection accepts both protein and nucleic acid residues, including mixed structures. |
| HETATM residues | enum | 0 (Ignore HETATM (default)) | 1: Include HETATM residues; 2: Include MSE residues only, alongside ATOM residues. Mirror comparisons recommend including eligible HETATM residues for D amino acids, but preserve the selected value. |
| Fast alignment | boolean | false | Uses fTM-align for faster, potentially less accurate alignment. |
| Alignment mode option | Value | Native behavior |
|---|---|---|
| Monomer (default) | monomer | Aligns monomeric structures or individual chain pairs. |
| Oligomer (multi-chain) | oligomer | Aligns multi-chain complexes. |
| Circular permutation | circular | Aligns structures with reordered termini. |
| Mirror image | mirror | Mirrors the mobile structure before monomer alignment. |
| Monomers to complex | monomers | Aligns a collection of monomers to a reference complex; requires Align monomers to a complex inputs. |
| Multiple structure alignment | multiple | Forms a consensus alignment from a collection; requires Multiple structure alignment inputs. |
| Fully non-sequential | nonsequential | Uses the native fully non-sequential algorithm. |
| Semi-non-sequential | semisequential | Uses the native semi-non-sequential algorithm. |
| Flexible alignment | flexible | Uses the native flexible-alignment algorithm with the selected hinge limit. |
Chain selection
| Parameter | Type | Default | Description |
|---|---|---|---|
| Structure 1 chains | string | optional (empty) | Comma-separated mobile chain IDs, such as A,B; _ represents a blank chain ID. Empty uses all eligible chains subject to Chain handling. |
| Structure 2 chains | string | optional (empty) | Comma-separated reference chain IDs with the same syntax and chain-handling restriction. |
Advanced native options
| Parameter | Type | Default | Description |
|---|---|---|---|
| Mirror the mobile structure | boolean | false | Enables native mirroring in other alignment modes; Mirror image alignment enables mirroring regardless of this switch. |
| Structure 1 models | string | optional (empty) | Comma-separated native mobile MODEL IDs, such as 1,2; empty uses the native default. Chain handling still governs which models are eligible. |
| Structure 2 models | string | optional (empty) | Comma-separated reference MODEL IDs with the same syntax and chain-handling restriction. |
| Representative atom | string | auto | Required, 1 to 4 printable ASCII characters; spaces are preserved. auto selects protein CA or nucleic acid C3'; PC4' selects the native hybrid P/C4' rule. |
| Split structures | enum | auto (Native default) | 0: Keep as one chain; 1: Split by MODEL; 2: Split by chain. The native default keeps one chain for residue-equivalence mode 2 and splits by chain otherwise. |
| Structure 1 format | enum | -1 (Auto-detect PDB/mmCIF) | 0: PDB; 1: SPICKER; 2: USAlign XYZ; 3: mmCIF. SPICKER and XYZ require an explicit choice. |
| Structure 2 format | enum | -1 (Auto-detect PDB/mmCIF) | Uses the same format choices for the reference input. |
| Residue equivalence | enum | 0 (Structural alignment) | Native residue correspondence modes 0 through 7 are listed below. |
| Sequence-dependent alignment (-seq) | boolean | false | Native alias for Glocal sequence alignment, residue-equivalence mode 5; Residue equivalence must be 0 or 5. |
| Evaluate existing superposition (-se) | boolean | false | Extracts an alignment without rotating or translating the input structures. |
| Nearest atoms for initial non-sequential alignment | integer | -1 | Minimum -1; that value preserves the native default of 5 for Fully non-sequential and 0 otherwise. A nonnegative value selects the native nearest-atom initialization count. |
| Maximum flexible hinges | integer | 9 | Range 0 to 9; applies to Flexible alignment. |
| Residue equivalence option | Value |
|---|---|
| Structural alignment | 0 |
| Same residue index | 1 |
| Same residue index and chain ID | 2 |
| Same residue index and chain order | 3 |
| Global sequence alignment | 4 |
| Glocal sequence alignment | 5 |
| Same residue index after chain mapping | 6 |
| Sequence alignment after chain mapping | 7 |
Score normalization
| Parameter | Type | Default | Description |
|---|---|---|---|
| Additional TM-score normalization | enum | native (Native default) | T: Average length; F: Second structure; -1: Shorter length; -2: Longer length; 1: Average length (numeric). Additional scores retain their native normalization labels. |
| Assigned normalization length (-u) | number | 0 | Minimum 0; 0 omits the option. A positive length adds a native score without changing the alignment. An assigned length shorter than the structure can produce a score above 1. |
| Assigned d0 in angstrom (-d) | number | 0 | Minimum 0; 0 omits the option. A positive distance adds the native distance-scaled score without changing the alignment. |
| TM-score early-stop cutoff | number | -1 | -1 disables early stopping. Native guidance is 0.5 to below 1; unlikely comparisons can stop early. The cutoff uses the selected additional normalization. |
Native outputs
| Parameter | Type | Default | Description |
|---|---|---|---|
| Coordinate and visualization files | enum | pymol (PyMOL) | chimerax: ChimeraX; rasmol: RasMol; none: Report only. Coordinates are produced only for supported input arrangements and formats. |
| Save rotation matrix | boolean | false | Saves the native rotation matrix when supported; list-comparison modes print matrices in the native report. |
| Report aligned residue distances | boolean | false | Adds the native aligned-residue distance table to the report. |
| Include individual alignments in complex or multiple results | boolean | false | Requests native full individual alignments, particularly for Monomers to complex and Multiple structure alignment; requires a multimer alignment mode. |
| Native report format | enum | 0 (Full report) | 1: Compact FASTA; 2: Compact tabular; -1: Full report without citation header. Compact tabular restricts additional normalization as described below. |
Compatible combinations
Incompatible settings are rejected rather than silently changed:
- Oligomer (multi-chain) and Monomers to complex require Chain handling
auto,0or1. Residue equivalence2,3and6have the same chain-handling requirement. - Use a chain mapping requires Oligomer (multi-chain). Use a supplied alignment requires Monomer (default) or Mirror image, Residue equivalence
0, and Sequence-dependent alignment (-seq) off. - Align monomers to a complex and Multiple structure alignment inputs require their matching alignment modes, and those modes require the matching input arrangements.
- Oligomer (multi-chain), Monomers to complex, Multiple structure alignment, Fully non-sequential, Semi-non-sequential and Flexible alignment require Assigned normalization length (-u)
0, Residue equivalence0, and Sequence-dependent alignment (-seq) off. - Compact tabular requires Additional TM-score normalization
nativeorF, Assigned normalization length (-u)0, and Assigned d0 in angstrom (-d)0. - Include individual alignments in complex or multiple results requires Oligomer (multi-chain), Monomers to complex, Multiple structure alignment, Fully non-sequential, Semi-non-sequential or Flexible alignment.
Job details
| Parameter | Type | Default | Description |
|---|---|---|---|
| Job name | string | optional (empty) | Labels the job in its history; it does not change the scientific calculation. |
Scores and files
The data table retains every native alignment record in source order, including structure/chain identifiers and every reported TM-score normalization. Sequence identity is a fraction, such as 0.25 for 25%. RMSD is in angstrom. A score normalized by the fixed reference length is useful when comparing a prediction with an experimental reference. Always compare scores using the same normalization.
Additional normalization options report scores using the average, shorter or longer structure length, an assigned length (-u), or an assigned distance scale in angstrom (-d). These options do not change the final native alignment. An assigned length shorter than the structure can produce a TM-score above 1. The native TM-score cutoff can stop unlikely comparisons early.
Choose PyMOL, ChimeraX, RasMol or report-only output. Downloads retain the native coordinate variants and original scripts, a complete native report, separate stdout, stderr when present, and JSON provenance with the command, source revision, captured-input hashes and native-file hashes. PyMOL and ChimeraX scripts also have additional portable copies that reference download filenames. Enable the rotation matrix or aligned-residue distance table when needed. Compact FASTA and tabular reports preserve their native text.
When monomer alignment processes several chain pairs, USAlign writes the superposition for the last pair. Coordinate metadata identifies that record, while the table retains all pairs. Multi-structure or complex reports can have several records without a single per-record coordinate association; the complete native report remains available.
Workflows provide the aligned structure, metric table and a collection of reports, scripts and supporting files, including rotation matrices when produced. A scoring-only native operation explicitly omits the aligned structure.
Source
See the USAlign source and native command-line documentation for file syntax and algorithm details. Historical jobs retain their admitted execution behavior; their stored settings and provenance identify that behavior.
Related tools
MAFFT
Align protein or nucleotide sequences with selectable accuracy and speed trade-offs.
MUSCLE5
Align multiple protein or nucleotide sequences with high-accuracy PPP refinement.
StringZilla v5
Hardware-accelerated edit distances and global or local sequence scores
Clustal Omega
Align multiple protein or nucleotide sequences and export FASTA, Clustal, or Phylip outputs.
FastTree
Build phylogenetic trees from aligned protein or nucleotide sequences using approximate maximum-likelihood methods.
FoldSeek
Search AlphaFold DB, compare structures, or cluster by 3D similarity
IQ-TREE
Build maximum likelihood phylogenetic trees with automatic model selection and ultrafast bootstrap.
MMseqs2
Search and cluster protein or nucleotide sequences for homology discovery at large scale.
MUMmer4
Align and compare whole genomes to detect SNPs, indels, and structural variants.
DockQ
Assess docking quality between model and native structures