What is protein design?

Protein design is the computational creation or optimization of protein sequences and structures for a defined objective. The objective may be a new fold, a sequence compatible with a backbone, catalytic activity, antigen recognition, peptide binding, target-specific protein binding, or improved sequence properties. Different design tasks require different inputs, models, scores, and experimental evidence.

Protein design is an umbrella use case rather than one algorithm. De novo methods generate new structural or sequence space; inverse folding finds sequences for a fixed backbone; enzyme, antibody, peptide, and binder design add functional and molecular-class constraints. Protein sequence design spans both structure-conditioned and sequence-generative approaches.

Choose the narrowest task that matches what you are trying to create. Every workflow should preserve constraints, random seeds, generated candidates, component scores, structures, and rejection reasons. Computational ranking reduces a search space, while synthesis, expression, structural characterization, and functional assays determine whether a design works.

When to use protein design

  • Create a new protein candidate. Generate a backbone, sequence, binder, enzyme, antibody, or peptide against an explicit objective.
  • Redesign an existing structure. Search sequences compatible with a backbone while preserving fixed functional or structural positions.
  • Connect design with validation. Keep generation, sequence assignment, refolding, property review, and experimental handoff in one inspectable record.

Benefits of protein design

  • Broader search. Explore structural and sequence possibilities beyond manual mutation.
  • Explicit constraints. Encode geometry, residues, targets, and properties as inspectable design requirements.
  • Connected evidence. Carry candidates from generation through model review and experimental handoff.

Primary limitations

  • Model-dependent rankings. Scores are approximations tied to each method and training distribution.
  • Incomplete biological context. Dynamics, expression systems, partners, and cellular conditions may be missing.
  • Experiments remain decisive. Computational candidates require synthesis, characterization, and functional testing.

Types of protein design

These seven use cases represent distinct searched tasks with different inputs, outputs, constraints, and validation workflows.

De novo protein design

Generates new protein backbones and sequences rather than modifying a supplied natural template.

Best for: Creating new folds, assemblies, or functional scaffolds
Requires: A design objective, constraints, and a validation plan

Inverse folding

Searches for amino-acid sequences expected to adopt a supplied three-dimensional backbone.

Best for: Fixed-backbone redesign and sequence recovery
Requires: A clean protein backbone structure

Enzyme design

Designs catalytic scaffolds and ligand-aware sequences around active-site geometry.

Best for: New or altered catalytic activity
Requires: Catalytic geometry, ligand or cofactor context, and an assay

Antibody design

Generates or redesigns antibody and nanobody sequences, structures, and binding loops.

Best for: Antigen-specific biologic discovery and optimization
Requires: An antigen, framework context, or antibody structure

Peptide design

Generates short peptide sequences for binding or other desired molecular properties.

Best for: Compact binders and peptide therapeutic leads
Requires: A target or property objective and length constraints

Protein sequence design

Creates or optimizes amino-acid sequences against structural, functional, or developability goals.

Best for: Sequence generation, redesign, and multi-objective optimization
Requires: A defined objective and optional structural context

Protein binder design

Designs proteins intended to recognize a specified target surface or epitope.

Best for: Target-specific research binders and therapeutic starting points
Requires: A target structure and a defined binding surface

Choosing a protein design method

The useful taxonomy follows researcher intent. De novo protein design creates new structures; inverse folding solves sequence for a fixed backbone; enzyme, antibody, peptide, and protein-binder design add distinct functional or molecular constraints; protein sequence design covers broader sequence generation and optimization.

Ligand-conditioned design and stability design remain important methods and objectives, but they fit within enzyme, sequence, or binder workflows here rather than becoming additional maintained spokes without demonstrated search demand.

How to design proteins online

A defensible protein-design workflow begins with a measurable objective and ends with experimental evidence. The software stages should make every assumption, constraint, candidate, and acceptance gate visible.

  1. Define the objective. State the desired structure, interaction, reaction, or property and how success will be measured.
  2. Prepare constraints. Confirm structures, sequences, motifs, fixed residues, targets, and allowed design space.
  3. Generate candidates. Use a method whose documented task and input requirements match the project.
  4. Evaluate evidence. Review native scores, refolding, geometry, diversity, stability, solubility, and task-specific evidence.
  5. Test experimentally. Select a diverse panel and measure expression, structure, and the intended function.

How to evaluate protein designs

No single score establishes design success. Evaluate constraint satisfaction, structural agreement, local confidence, geometry, sequence diversity, stability, solubility, aggregation, and task-specific evidence separately.

Compare candidates within the same model and settings unless scores are explicitly calibrated across methods. Preserve rejected designs and score distributions to avoid selection based only on attractive visualizations.

Experimental validation for protein design

Expression and folding are early gates, not proof of function. The decisive assay must match the objective: affinity and specificity for binders, catalytic rate and selectivity for enzymes, or structural and biophysical properties for scaffold designs.

Plan controls, replication, and candidate diversity before generation. Export sequences, structures, settings, provenance, and selection logic together for synthesis and laboratory handoff.

How protein design works

The hub workflow compares four sequence-design methods from one backbone; each spoke uses a workflow matched to its specific design task.

  1. Define the objective. Choose the protein-design task and acceptance criteria.
  2. Prepare the backbone. Clean the PDB and identify chains and fixed residues.
  3. Run four methods. Submit the same backbone to four sequence-design models.
  4. Compare candidates. Review probabilities, diversity, conserved positions, and method agreement.
  5. Validate and export. Refold and screen selected candidates before experimental testing.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Design evidence. PDB FASTA JSON TXT Protein structures, sequences, targets, motifs, fixed residues, and constraints required by the selected design task.

Outputs

  • Design candidates. PDB FASTA CSV JSON Generated structures and sequences, model-native scores, rankings, logs, and validation files.

Featured protein design workflow

Preserve four complementary sequence-design outputs for direct comparison and downstream validation.

Protein design method panelRead-only preview

Inputs

1 required

Methods

4 connected

  1. 01ProteinMPNN
  2. 02ESM-IF1
  3. 03SolubleMPNN
  4. 04HyperMPNN

Preserve four complementary sequence-design outputs for direct comparison and downstream validation.

Use this template

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open comparison workflow