Use case

Protein sequence design

Generate or optimize amino-acid sequences against structural, functional, stability, and developability objectives.

Protein design -> stability screenRead-only preview

Inputs

1 required

Methods

6 connected

  1. 01ProteinMPNN
  2. 02ESMFold
  3. 03Protein Stability
  4. 04NetSolP-1.0
  5. 05Aggrescan3D
  6. 06SASA Calculator

Redesign a backbone with ProteinMPNN, refold sequences, and screen stability, solubility, aggregation, and surface exposure.

Use this template

What is protein sequence design?

Protein sequence design is the computational process of finding amino-acid sequences expected to satisfy one or more structural or functional objectives. Some methods condition on a fixed backbone, while protein language models can generate sequences with weaker or no explicit structural input. Projects may optimize folding, stability, solubility, expression, binding, catalysis, or combinations of these properties, but each objective needs its own evidence and validation.

Sequence design is broader than inverse folding. Inverse folding specifically asks which sequences fit a supplied backbone; sequence-design projects may instead generate from a language model, edit an existing protein, or combine structural and property objectives in an iterative search.

Define immutable residues, allowed mutations, diversity targets, and acceptance gates before sampling. Multi-objective design should retain the separate component scores because a single aggregate rank can hide a severe failure in stability, function, or developability.

When to use protein sequence design

  • Best fit. Generating new sequences or redesigning existing proteins against explicit objectives
  • Required starting evidence. A sequence or backbone context, mutable positions, objective definitions, and validation assays

Benefits of protein sequence design

  • Focused search. Explores mutations jointly
  • Connected evidence. Supports explicit residue constraints
  • Testable candidates. Can balance several design objectives

Primary limitations

  • Model scope. Sequence space is enormous
  • Score uncertainty. Objectives may conflict
  • Experimental requirement. Predictive scores may be poorly calibrated

Protein sequence design methods

Backbone-conditioned models use local and global geometry to propose residues, while protein language models learn statistical regularities from natural sequence data. These approaches can be complementary but their scores do not share a universal scale.

Optimization algorithms can search several objectives, yet the weighting scheme encodes project priorities. Hard requirements should be treated as gates where possible rather than diluted inside one weighted average.

How to run protein sequence design online

Use the workflow as an inspectable computational funnel. Preserve the native output of each method, apply explicit acceptance gates, and keep the evidence behind every selected and rejected candidate.

  1. Define objectives. Define the structural, functional, and developability objectives and how each will be evaluated.
  2. Set constraints. Specify fixed residues, mutable regions, sequence identity limits, and forbidden motifs.
  3. Generate sequences. Generate a diverse sequence set with a backbone-conditioned or sequence-generative method.
  4. Screen properties. Evaluate refolding, stability, solubility, aggregation, and task-specific function without hiding component scores.
  5. Test candidates. Select a diverse experimental panel and measure expression, structure, and the intended function.

How to evaluate protein sequence design results

Review diversity, similarity to training or parent sequences, conserved functional residues, refold agreement, stability, solubility, aggregation, and any task-specific predictive endpoint.

Experimental selection should cover more than the numerical top rank. A diverse panel is more informative about which design assumptions generalize and can reveal score failures.

Experimental validation and handoff

Keep all objective definitions, fixed and mutable positions, sampling settings, component scores, and rejected sequences. Test structure and intended function experimentally.

Export structures, sequences, settings, scores, logs, and selection criteria together. A reproducible handoff makes computational assumptions visible to the team planning synthesis, expression, biophysical characterization, and functional assays.

How protein sequence design works

Redesign a backbone with ProteinMPNN, refold sequences, and screen stability, solubility, aggregation, and surface exposure.

  1. Define objectives. Define the structural, functional, and developability objectives and how each will be evaluated.
  2. Set constraints. Specify fixed residues, mutable regions, sequence identity limits, and forbidden motifs.
  3. Generate sequences. Generate a diverse sequence set with a backbone-conditioned or sequence-generative method.
  4. Screen properties. Evaluate refolding, stability, solubility, aggregation, and task-specific function without hiding component scores.
  5. Test candidates. Select a diverse experimental panel and measure expression, structure, and the intended function.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Design input. PDB FASTA JSON TXT A protein backbone or starting sequence plus fixed positions, mutable regions, and design objectives.

Outputs

  • Design and review outputs. PDB FASTA CSV JSON Designed FASTA sequences, per-objective scores, refolded models, property tables, and candidate rankings.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow