Use case
De novo protein design
Generate proteins beyond known templates, then connect backbone generation, sequence assignment, refolding, and candidate review.
Inputs
0 required
Methods
4 connected
- 01RFdiffusion3 · Unconditional Design
- 02ProteinMPNN
- 03ESMfold · Refold Check
- 04MolProbity
Generate unconditional backbones with RFdiffusion3, assign sequences with ProteinMPNN, refold with ESMfold, and inspect geometry with MolProbity.
Use this templateWhat is de novo protein design?
De novo protein design is the process of creating new protein structures and sequences from scratch rather than modifying a natural protein that already performs the intended role. Generative models can propose backbones under geometric or functional constraints, after which sequence-design methods assign residues expected to stabilize those structures. The result is a set of computational hypotheses that still requires structural and experimental validation.
A complete campaign separates backbone generation from sequence assignment and validation. This matters because a visually plausible backbone does not guarantee that any amino-acid sequence will fold into it, and a high model score does not establish soluble expression, monodispersity, stability, or function.
Use unconditional generation for new structural space, or add symmetry, motif, target, and functional constraints when the design must satisfy a specific geometry. Preserve rejected candidates and score distributions so selection is based on stated gates rather than one attractive model.
When to use de novo protein design
- Best fit. New folds, assemblies, functional scaffolds, and constrained structural concepts
- Required starting evidence. A design objective, structural constraints, candidate count, and experimental acceptance criteria
Benefits of de novo protein design
- Focused search. Explores structure beyond natural templates
- Connected evidence. Supports explicit geometric constraints
- Testable candidates. Produces diverse testable hypotheses
Primary limitations
- Model scope. Computational success does not imply expression
- Score uncertainty. Scores are model-dependent
- Experimental requirement. Experimental hit rates may be low
De novo protein design methods
Diffusion and other generative models sample structures from learned protein geometry. Conditioning can steer that sampling toward a motif, target, symmetry, or shape, but stronger constraints can reduce diversity or create incompatible requirements.
Sequence assignment is a separate inverse problem. Multiple sequences per backbone help reveal whether a structural proposal has a broad compatible sequence space or depends on a narrow, fragile solution.
How to run de novo protein design online
Use the workflow as an inspectable computational funnel. Preserve the native output of each method, apply explicit acceptance gates, and keep the evidence behind every selected and rejected candidate.
- Specify intent. Define the desired size, topology, symmetry, motif, or functional constraints before generation.
- Generate backbones. Generate a sufficiently diverse backbone set and retain model settings and random seeds.
- Assign sequences. Assign several sequences to each accepted backbone instead of treating one sequence as definitive.
- Refold candidates. Refold sequences independently and compare predicted structures with their design backbones.
- Select for testing. Apply geometry, stability, solubility, and experiment-specific gates before synthesis.
How to evaluate de novo protein design results
Compare designed and independently predicted structures using global and local agreement, then inspect clashes, secondary structure, buried polar atoms, exposed hydrophobics, aggregation risk, and sequence diversity.
Prospective validation should measure the property the design was intended to create. Expression and folding checks are necessary, but they do not replace binding, catalytic, assembly, or other functional assays.
Experimental validation and handoff
Retain generation settings, seeds, every sequence–backbone pairing, refolded structures, and rejection criteria. Confirm folding and intended function experimentally.
Export structures, sequences, settings, scores, logs, and selection criteria together. A reproducible handoff makes computational assumptions visible to the team planning synthesis, expression, biophysical characterization, and functional assays.
How de novo protein design works
Generate unconditional backbones with RFdiffusion3, assign sequences with ProteinMPNN, refold with ESMfold, and inspect geometry with MolProbity.
- Specify intent. Define the desired size, topology, symmetry, motif, or functional constraints before generation.
- Generate backbones. Generate a sufficiently diverse backbone set and retain model settings and random seeds.
- Assign sequences. Assign several sequences to each accepted backbone instead of treating one sequence as definitive.
- Refold candidates. Refold sequences independently and compare predicted structures with their design backbones.
- Select for testing. Apply geometry, stability, solubility, and experiment-specific gates before synthesis.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Design input.
PDBFASTAJSONTXTDesign constraints and, when applicable, a target, motif, or symmetry definition.
Outputs
- Design and review outputs.
PDBFASTACSVJSONDesigned backbone PDB files, candidate FASTA sequences, refolded models, scores, and review files.
Tools for de novo protein design
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

RFdiffusion3
Generate all-atom protein backbones

RFdiffusion 2
Generate motif- and ligand-conditioned scaffolds

ProGen2
Generate protein sequences with a language model

EvoDiff
Generate sequence candidates with diffusion

ProteinMPNN
Assign sequences to generated backbones

ESM-IF1
Design alternative sequences for fixed backbones

ESMfold
Refold designed sequences for structural comparison

USAlign
Compare refolded models with design backbones

MolProbity
Review clashes and stereochemical geometry

Protein stability
Estimate sequence-level stability signals

NetSolP-1.0
Estimate sequence-level solubility

Aggrescan3D
Inspect structure-based aggregation-prone regions
Other protein engineering workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Inverse folding
Searches for amino-acid sequences expected to adopt a supplied three-dimensional backbone.
Enzyme design
Designs catalytic scaffolds and ligand-aware sequences around active-site geometry.
Antibody design
Generates or redesigns antibody and nanobody sequences, structures, and binding loops.
Peptide design
Generates short peptide sequences for binding or other desired molecular properties.
Protein sequence design
Creates or optimizes amino-acid sequences against structural, functional, or developability goals.
Protein binder design
Designs proteins intended to recognize a specified target surface or epitope.
Frequently asked questions
Use the same size range, structural constraints, candidate budget, sequence-design settings, and acceptance gates. Record method versions and random seeds so differences can be attributed to the generators rather than to different campaign setups.
Cluster both structures and sequences before final ranking, then cap the number selected from any one cluster. Increasing random seeds or sampling temperature can add diversity, but the resulting candidates still need to pass the same structural gates.
Select across distinct backbone and sequence clusters rather than taking only the numerical top ranks. A diverse panel tests more of the model’s assumptions and produces more useful evidence for the next design round.
Yes. Treat expression, solubility, oligomeric state, and folding outcomes as labeled campaign data. They can guide tighter filters or objectives, although a small project dataset is usually insufficient to retrain a generative model reliably.
Complete de novo protein-design programs that include sequence design, protein production, and experimental testing are quote-based among the current providers reviewed. Schrödinger quotes comprehensive protein-design services, while WuXi Biologics quotes 4–5 week miniprotein production and binding-assay projects; neither publishes one fixed design-to-validation package price.
The quote depends on the number of designs advanced, expression host, purification requirements, structural characterization, binding or functional assay, optimization rounds, and whether the provider is responsible for both computational design and laboratory validation.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured run quoted in credits before submission. A done-for-you de novo protein design project is scoped separately; synthesis, expression, and experimental assays are included only when the project quote explicitly says so.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.