Use case
Protein sequence design
Generate or optimize amino-acid sequences against structural, functional, stability, and developability objectives.
Inputs
1 required
Methods
6 connected
- 01ProteinMPNN
- 02ESMFold
- 03Protein Stability
- 04NetSolP-1.0
- 05Aggrescan3D
- 06SASA Calculator
Redesign a backbone with ProteinMPNN, refold sequences, and screen stability, solubility, aggregation, and surface exposure.
Use this templateWhat is protein sequence design?
Protein sequence design is the computational process of finding amino-acid sequences expected to satisfy one or more structural or functional objectives. Some methods condition on a fixed backbone, while protein language models can generate sequences with weaker or no explicit structural input. Projects may optimize folding, stability, solubility, expression, binding, catalysis, or combinations of these properties, but each objective needs its own evidence and validation.
Sequence design is broader than inverse folding. Inverse folding specifically asks which sequences fit a supplied backbone; sequence-design projects may instead generate from a language model, edit an existing protein, or combine structural and property objectives in an iterative search.
Define immutable residues, allowed mutations, diversity targets, and acceptance gates before sampling. Multi-objective design should retain the separate component scores because a single aggregate rank can hide a severe failure in stability, function, or developability.
When to use protein sequence design
- Best fit. Generating new sequences or redesigning existing proteins against explicit objectives
- Required starting evidence. A sequence or backbone context, mutable positions, objective definitions, and validation assays
Benefits of protein sequence design
- Focused search. Explores mutations jointly
- Connected evidence. Supports explicit residue constraints
- Testable candidates. Can balance several design objectives
Primary limitations
- Model scope. Sequence space is enormous
- Score uncertainty. Objectives may conflict
- Experimental requirement. Predictive scores may be poorly calibrated
Protein sequence design methods
Backbone-conditioned models use local and global geometry to propose residues, while protein language models learn statistical regularities from natural sequence data. These approaches can be complementary but their scores do not share a universal scale.
Optimization algorithms can search several objectives, yet the weighting scheme encodes project priorities. Hard requirements should be treated as gates where possible rather than diluted inside one weighted average.
How to run protein sequence design online
Use the workflow as an inspectable computational funnel. Preserve the native output of each method, apply explicit acceptance gates, and keep the evidence behind every selected and rejected candidate.
- Define objectives. Define the structural, functional, and developability objectives and how each will be evaluated.
- Set constraints. Specify fixed residues, mutable regions, sequence identity limits, and forbidden motifs.
- Generate sequences. Generate a diverse sequence set with a backbone-conditioned or sequence-generative method.
- Screen properties. Evaluate refolding, stability, solubility, aggregation, and task-specific function without hiding component scores.
- Test candidates. Select a diverse experimental panel and measure expression, structure, and the intended function.
How to evaluate protein sequence design results
Review diversity, similarity to training or parent sequences, conserved functional residues, refold agreement, stability, solubility, aggregation, and any task-specific predictive endpoint.
Experimental selection should cover more than the numerical top rank. A diverse panel is more informative about which design assumptions generalize and can reveal score failures.
Experimental validation and handoff
Keep all objective definitions, fixed and mutable positions, sampling settings, component scores, and rejected sequences. Test structure and intended function experimentally.
Export structures, sequences, settings, scores, logs, and selection criteria together. A reproducible handoff makes computational assumptions visible to the team planning synthesis, expression, biophysical characterization, and functional assays.
How protein sequence design works
Redesign a backbone with ProteinMPNN, refold sequences, and screen stability, solubility, aggregation, and surface exposure.
- Define objectives. Define the structural, functional, and developability objectives and how each will be evaluated.
- Set constraints. Specify fixed residues, mutable regions, sequence identity limits, and forbidden motifs.
- Generate sequences. Generate a diverse sequence set with a backbone-conditioned or sequence-generative method.
- Screen properties. Evaluate refolding, stability, solubility, aggregation, and task-specific function without hiding component scores.
- Test candidates. Select a diverse experimental panel and measure expression, structure, and the intended function.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Design input.
PDBFASTAJSONTXTA protein backbone or starting sequence plus fixed positions, mutable regions, and design objectives.
Outputs
- Design and review outputs.
PDBFASTACSVJSONDesigned FASTA sequences, per-objective scores, refolded models, property tables, and candidate rankings.
Tools for protein sequence design
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

ProteinMPNN
Design sequences for a supplied backbone

ESM-IF1
Generate geometry-conditioned sequence alternatives

ProGen2
Generate protein sequences with a language model

EvoDiff
Generate diverse protein sequences by diffusion

SolubleMPNN
Optimize sequences for soluble proteins

HyperMPNN
Bias sequence design toward thermostability

ESMfold
Refold designed sequences for structural comparison

Protein stability
Estimate sequence-level stability signals

NetSolP-1.0
Estimate sequence-level solubility

Aggrescan3D
Inspect structure-based aggregation-prone regions

SASA calculator
Measure solvent-accessible surface area

Protein parameters
Calculate sequence physicochemical properties
Other protein engineering workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
De novo protein design
Generates new protein backbones and sequences rather than modifying a supplied natural template.
Inverse folding
Searches for amino-acid sequences expected to adopt a supplied three-dimensional backbone.
Enzyme design
Designs catalytic scaffolds and ligand-aware sequences around active-site geometry.
Antibody design
Generates or redesigns antibody and nanobody sequences, structures, and binding loops.
Peptide design
Generates short peptide sequences for binding or other desired molecular properties.
Protein binder design
Designs proteins intended to recognize a specified target surface or epitope.
Frequently asked questions
Use a backbone-conditioned model when maintaining a specific three-dimensional geometry is central. Use a language model when broader sequence generation is acceptable, then add structure and property checks appropriate to the objective.
Use hard constraints for non-negotiable chemistry, fixed functional residues, forbidden motifs, and manufacturing rules. Reserve weighted objectives for genuine tradeoffs where partial improvement still has value.
Set the identity range from the project goal rather than from a universal cutoff. Conservative optimization may require close parent similarity, while scaffold diversification benefits from lower identity provided essential residues and structure remain supported.
Inspect every component score and the Pareto tradeoff rather than relying only on an aggregate rank. Include candidates that represent different defensible compromises so experiments can reveal which computational objective was most predictive.
Complete protein-sequence-design projects that include candidate generation, construct preparation, expression, purification, and functional testing are quote-based. Current protein-design and engineering service providers describe those project stages but do not publish one fixed sequence-to-validated-protein package price.
The quote depends on the number and diversity of sequences, optimization rounds, gene synthesis and cloning, expression host, screening scale, purification, biophysical characterization, and the functional assay required for the design objective.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured run quoted in credits before submission. A done-for-you protein sequence design project is scoped separately; synthesis, expression, and experimental assays are included only when the project quote explicitly says so.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.