Use case
Inverse folding
Design amino-acid sequences for a fixed protein backbone, compare complementary models, and refold candidates before selection.
Inputs
1 required
Methods
4 connected
- 01ProteinMPNN
- 02ESM-IF1
- 03ESMfold · ProteinMPNN Check
- 04ESMfold · ESM-IF1 Check
Design sequences from the same backbone with ProteinMPNN and ESM-IF1, then refold both branches independently.
Use this templateWhat is inverse folding?
Inverse folding is the task of finding amino-acid sequences that are compatible with a supplied three-dimensional protein backbone. Unlike structure prediction, which maps sequence to structure, inverse folding holds the backbone geometry fixed and predicts residue identities or sequence probabilities. It is widely used for fixed-backbone redesign, sequence recovery, stability-oriented diversification, and sequence assignment after backbone generation.
The backbone defines much of the structural context, but it does not uniquely determine a sequence. ProteinMPNN, ESM-IF1, and related models can return different high-probability solutions because they use different representations, training data, and decoding procedures.
Generate multiple sequences at controlled sampling temperatures, preserve per-residue probabilities, and avoid overinterpreting native-sequence recovery as proof of design quality. Refolding and developability filters can reject obvious failures, while experiments establish whether candidates actually adopt the intended structure.
When to use inverse folding
- Best fit. Fixed-backbone redesign, sequence recovery, and assigning sequences to designed structures
- Required starting evidence. A clean PDB backbone with intended chains, residues, and any fixed positions identified
Benefits of inverse folding
- Focused search. Directly conditions on 3D geometry
- Connected evidence. Samples many sequences per backbone
- Testable candidates. Supports residue-level constraints
Primary limitations
- Model scope. Treats the backbone as largely fixed
- Score uncertainty. Model probabilities are not experimental fitness
- Experimental requirement. Missing context can mislead design
Inverse folding methods
Inverse-folding networks encode backbone geometry and estimate compatible residue identities. Autoregressive and masked approaches differ in how sequence positions condition one another, so model comparison can expose candidates that depend on a single scoring assumption.
Sampling temperature controls the tradeoff between high-probability residues and sequence diversity. Fixed-position masks should preserve residues whose chemistry or interactions are essential rather than asking the model to rediscover every constraint.
How to run inverse folding online
Use the workflow as an inspectable computational funnel. Preserve the native output of each method, apply explicit acceptance gates, and keep the evidence behind every selected and rejected candidate.
- Prepare backbone. Clean the backbone and confirm chain identities, residue numbering, missing atoms, and intended oligomeric context.
- Define fixed residues. Specify residues that must remain fixed, including catalytic, binding, disulfide, or interface positions.
- Sample sequences. Run complementary inverse-folding models and sample multiple sequences at documented temperatures.
- Refold designs. Refold candidates and compare local and global agreement with the supplied backbone.
- Rank candidates. Review diversity, stability, solubility, and experimental constraints before selecting sequences.
How to evaluate inverse folding results
Review sequence log-probabilities, diversity, conserved positions, refold agreement, clashes, and local confidence. Compare candidates within the same model and settings because raw scores are not necessarily calibrated across methods.
A fixed-backbone design can fail when the real protein relaxes, changes oligomeric state, binds a partner, or encounters cellular constraints. Test structure, stability, and the intended function with appropriate experiments.
Experimental validation and handoff
Keep the exact backbone, fixed-position masks, sampling temperature, model version, sequence scores, and refolded models. Validate selected designs experimentally.
Export structures, sequences, settings, scores, logs, and selection criteria together. A reproducible handoff makes computational assumptions visible to the team planning synthesis, expression, biophysical characterization, and functional assays.
How inverse folding works
Design sequences from the same backbone with ProteinMPNN and ESM-IF1, then refold both branches independently.
- Prepare backbone. Clean the backbone and confirm chain identities, residue numbering, missing atoms, and intended oligomeric context.
- Define fixed residues. Specify residues that must remain fixed, including catalytic, binding, disulfide, or interface positions.
- Sample sequences. Run complementary inverse-folding models and sample multiple sequences at documented temperatures.
- Refold designs. Refold candidates and compare local and global agreement with the supplied backbone.
- Rank candidates. Review diversity, stability, solubility, and experimental constraints before selecting sequences.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Design input.
PDBFASTAJSONTXTA protein backbone in PDB format plus optional fixed-position and chain-design constraints.
Outputs
- Design and review outputs.
PDBFASTACSVJSONDesigned FASTA sequences, model scores, sequence alignments, refolded PDB files, and comparison results.
Tools for inverse folding
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

ProteinMPNN
General fixed-backbone sequence design

ESM-IF1
Geometric language-model inverse folding

SolubleMPNN
Sequence design specialized for soluble proteins

HyperMPNN
Sequence design biased toward thermostability

LigandMPNN
Ligand-aware inverse folding

AntiFold
Antibody-specialized inverse folding

ESMfold
Refold designed sequences for structural comparison

USAlign
Compare refolded candidates with the input backbone

MolProbity
Review clashes and stereochemical geometry

Protein stability
Estimate sequence-level stability signals

NetSolP-1.0
Estimate sequence-level solubility

Aggrescan3D
Inspect structure-based aggregation-prone regions
Other protein engineering workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
De novo protein design
Generates new protein backbones and sequences rather than modifying a supplied natural template.
Enzyme design
Designs catalytic scaffolds and ligand-aware sequences around active-site geometry.
Antibody design
Generates or redesigns antibody and nanobody sequences, structures, and binding loops.
Peptide design
Generates short peptide sequences for binding or other desired molecular properties.
Protein sequence design
Creates or optimizes amino-acid sequences against structural, functional, or developability goals.
Protein binder design
Designs proteins intended to recognize a specified target surface or epitope.
Frequently asked questions
Start near the method’s documented default and generate a small pilot at lower and higher temperatures. Lower values favor high-probability residues; higher values increase diversity and usually require stronger downstream filtering.
Repair only regions supported by defensible structural evidence, or exclude them from design. An invented loop becomes a design constraint, so its uncertainty should be explicit rather than silently treated as experimental geometry.
Not as one calibrated scale. Rank candidates within each model and sampling setup, then compare shared downstream evidence such as refold agreement, sequence diversity, and property checks.
Use the assembly that contains the interfaces the sequence must support. Designing an isolated monomer can expose or mutate residues that are buried in the biological oligomer, producing candidates incompatible with the intended complex.
Complete inverse-folding projects that carry designed sequences through gene construction, expression, purification, and structural or functional testing are quote-based. The protein-design and protein-engineering providers reviewed publish their scope but not a fixed end-to-end inverse-folding package price.
The total changes with backbone count, sequences per backbone, fixed-position constraints, gene and construct preparation, expression screening, purification, structural confirmation, and the assay used to decide whether a sequence works.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured run quoted in credits before submission. A done-for-you inverse folding project is scoped separately; synthesis, expression, and experimental assays are included only when the project quote explicitly says so.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.