Structure analysis
Protein structure search
Search large protein-structure databases with a query fold, then inspect ranked neighbors, coverage, scores, and alignments.
What is protein structure search?
Protein structure search is the process of querying a structural database with a three-dimensional protein model to find geometrically similar entries. FoldSeek represents local tertiary interactions with a structural alphabet, enabling fast candidate retrieval before detailed score, alignment, and biological review.
Use it to find remote-homolog candidates, fold analogs, related domains, or structural neighbors when sequence search is insufficient. The query may be experimental or predicted, but chain choice, domain boundaries, missing residues, and model confidence affect retrieval.
ProteinIQ runs FoldSeek database search directly against supported collections. Results retain identifiers, alignment statistics, TM-score and LDDT context, coverage, E-values, and files for detailed pairwise and biological review.
When to use protein structure search
- Structural neighbors are needed. Use it for remote homology, analogs, and fold-family context.
- A reviewed query model is available. Choose an appropriate chain and domain boundary.
- Hits can be independently assessed. Add sequence, annotation, and experimental evidence for important candidates.
Benefits of protein structure search
- Remote relationships can be found. It can reveal similarities missed by sequence search.
- Large databases can be searched. Retrieval scales to structural collections.
- Native evidence remains inspectable. Alignments and method-specific scores are returned.
Primary limitations
- Coverage bounds discovery. Absent database entries cannot be found.
- Query quality changes ranking. Boundaries and coordinate quality matter.
- Similarity does not prove function. Functional transfer needs independent evidence.
Protein structure search methods and applications
FoldSeek converts tertiary residue neighborhoods into a structural alphabet and applies sequence-search techniques. Database composition also matters: clustered predicted structures can reveal broad fold neighborhoods, while curated experimental entries may carry stronger ligand, assembly, and functional context.
Structure search supports remote-homology discovery, annotation, fold classification, model-quality investigation, target comparison, and template selection. A match may represent a global fold, shared domain, repeat, or local geometry, so inspect the aligned region before transferring database annotation.
How to run protein structure search online
- Prepare the query. Select chain, domain, assembly, and state; review missing or low-confidence regions.
- Choose databases. Set collections and thresholds for the needed sensitivity and result volume.
- Run FoldSeek. Preserve database version, settings, warnings, and failed inputs.
- Review matches. Read rank, E-value, coverage, TM-score, LDDT, alignment, and length differences together.
- Validate candidates. Confirm important hits with superposition, sequence evidence, annotations, and experiments.
How to interpret protein structure search results
Evaluate E-value, score, coverage, aligned length, TM-score, LDDT, and sequence identity together. Structural similarity alone does not prove common ancestry or biochemical function; confirm architecture, conserved residues, assembly, ligands, taxonomy, and sequence evidence.
How protein structure search works
FoldSeek performs the core structure search directly; downstream homology and function claims still require independent review.
- Prepare the query. Choose the relevant chain, domain, assembly, and conformational state, and review missing or low-confidence regions.
- Choose databases. Select structural databases and thresholds that match the intended sensitivity and result volume.
- Run FoldSeek. Run FoldSeek while preserving database versions, search settings, warnings, and failed inputs.
- Review matches. Inspect rank, E-value, coverage, TM-score, LDDT, residue alignment, and query–target length differences together.
- Validate candidates. Confirm important candidates with detailed superposition, sequence evidence, curated annotations, and experiments when the claim requires them.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
Structure-analysis inputs
PDBmmCIFFASTATSVOne experimental or predicted protein structure in PDB or mmCIF format.
Outputs
Reviewable results
PDBCSVTSVJSONFILESRanked database hits, identifiers, E-values, coverage, TM-scores, LDDT values, alignments, and downloadable files.
Tools for protein structure search
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

FoldSeek
Search structure databases or compare and cluster uploaded protein structures

USAlign
Align two protein structures and return TM-scores, RMSD, residue correspondence, and superposed coordinates

PDBFixer
Repair common coordinate-file issues before structural comparison

PDB Download
Retrieve experimental structures from the Protein Data Bank

AlphaFold Database Download
Retrieve predicted protein structures from the AlphaFold Protein Structure Database

PDB to FASTA converter
Extract protein sequences from coordinate files for sequence-aware review

HMMER
Search profile hidden Markov models for independent sequence-level homology evidence

MMseqs2
Search and cluster large protein sequence collections

DSSP
Assign secondary structure and solvent accessibility from protein coordinates

MolProbity
Check model geometry and steric quality before interpreting structural matches

SASA calculator
Calculate solvent-accessible surface area for matched structures

RMSD calculator
Superpose comparison structures on one reference and report RMSD values
Other structure analysis workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Frequently asked questions
One experimental or predicted protein structure in PDB or mmCIF format.
Ranked database hits, identifiers, E-values, coverage, TM-scores, LDDT values, alignments, and downloadable files.
Confirm accession, model, chain, biological assembly, domain boundaries, residue numbering, missing regions, alternate conformations, and prediction confidence. Repair coordinates only when necessary and retain both the original file and every preparation decision.
Use method-native scores together rather than selecting one universal number. TM-score emphasizes length-normalized global fold similarity, RMSD reports geometric deviation over the aligned atoms, and coverage shows how much of each structure actually corresponds.
No. Similar folds can support different functions, and local similarity can occur without shared global architecture. Review residue-level correspondence, domains, ligands, oligomeric state, taxonomy, sequence evidence, curated annotations, and experiments.
A complete protein structure search project is generally quote-based. Current providers describe fold recognition and protein-structure analysis as customized services covering data review, method selection, modeling or comparison, validation, and interpretation rather than publishing one universal project price.
The cost depends on structure or sequence count, database scope, model preparation, method comparison, manual inspection, figures, annotation, and whether experimental follow-up is included. Open-source FoldSeek, US-align, and FoldMason can remove a software-license fee, but they do not remove expert analysis or compute requirements.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured protein structure search workflow estimated in credits before submission. Done-for-you analysis is scoped separately.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.