
Design protein structures for de novo scaffolds, binders, motifs, and symmetric oligomers. Learn more
Input
What is RFdiffusion?
RFdiffusion is a generative AI model for designing novel protein structures with atomic precision. Developed by the Baker Lab at the University of Washington and published in Nature (2023), it uses diffusion models—the same technology behind DALL-E and Stable Diffusion—adapted for protein structure generation. With over 1,000 citations since publication, the method achieves approximately two orders of magnitude improvement over traditional computational protein design approaches.
Independent experimental validation demonstrates 84% confirmation rates for designed binders, with affinities ranging from tens of micromolar to tens of nanomolar. Applications include therapeutic binder design, enzyme active site scaffolding, symmetric oligomer creation, and de novo protein generation.
RFdiffusion works by iteratively denoising random 3D coordinates into structured protein backbones through learned reverse diffusion steps. Unlike traditional methods that optimize physics-based energy functions, it learns directly from the statistical distribution of natural protein structures in the Protein Data Bank.
How does RFdiffusion work?
Architecture
RFdiffusion builds on the RoseTTAFold (RF) structure prediction network architecture, similar to AlphaFold2. The model was initialized with RoseTTAFold's pretrained weights and fine-tuned specifically for structure denoising tasks. It uses SE(3)-equivariant graph neural networks that respect rotational and translational symmetries, improving generalization and data efficiency.
The architecture processes protein structures as rigid-frame representations—coordinates and orientations of backbone atoms (N, Cα, C, O, and virtual Cβ). This representation enables precise control at the residue level while maintaining physical realism.
Diffusion process
The model operates through reverse diffusion on protein backbone structures. Starting from random noise (random 3D coordinates), it iteratively denoises through learned steps toward a structured protein. At each timestep, the model predicts the final structure from the current noised structure, then interpolates from the current coordinates toward the predicted structure.
Training used 200 discrete diffusion timesteps, though the default deployment uses 50 steps for optimal quality-speed balance. Research shows that as few as 20 timesteps achieve equivalent quality with 10x speedup. Higher timestep counts improve quality with diminishing returns beyond 50.
Self-conditioning
A critical innovation is self-conditioning: the model receives its previous prediction as template input, similar to AlphaFold2's "recycling" mechanism. This architectural choice improves prediction quality by allowing the model to iteratively refine its understanding of the emerging structure.
Key innovations
Unlike traditional protein design methods requiring manual parameter tuning and physics-based optimization, RFdiffusion treats design as generative modeling over protein structure space. It requires no evolutionary information or multiple sequence alignments (MSAs). The approach enables diverse design tasks through conditional generation—specifying constraints like binding sites, motifs, or symmetry while allowing the model to generate compatible structures.
Design modes
RFdiffusion supports five distinct design modes for different applications:
Binder design
Creates proteins that bind to a specified target protein. You provide the target structure and optionally specify binding pocket residues (hotspots). The model generates novel binder proteins with specified length ranges that form interfaces with the target. This mode has been experimentally validated with picomolar-affinity binders to therapeutic targets including MDM2, PD-L1, and IL-7Rα.
Motif scaffolding
Builds a protein scaffold around a functional motif of interest. You specify which residues to preserve and how much structure to add at the N- and C-termini. Applications include scaffolding viral epitopes, receptor binding sites, enzyme active sites, and metal-binding motifs. This enables transplanting functional elements into new structural contexts.
Partial diffusion
Partially redesigns an existing protein structure to create variants. By controlling the diffusion timesteps (partial_T), you can tune the diversity—low values create subtle variations, high values generate more radical changes. Useful for exploring structural space around a known fold while maintaining key geometric features.
Unconditional generation
Generates novel protein structures from scratch without any template. You specify the desired length and optional symmetry. The model samples entirely new protein folds from learned structural distributions. This mode enables exploring uncharted regions of protein structure space.
Custom design
Advanced mode for users who want precise control using contig syntax. Contigs define exactly which regions to keep from input structures and where to generate new structure. This enables complex design tasks like inserting domains, creating fusion proteins, or specifying exact topological arrangements.
Input requirements
RFdiffusion accepts PDB format structures (.pdb, .ent files) up to 50 MB. You can upload local files or provide RCSB PDB IDs to fetch structures directly. Most modes require an input structure, except unconditional generation which starts from scratch.
Input parameters
- Number of designs: Controls how many independent structures to generate in a single job. Default: 10. Range: 1-50 (linear scaling with computation time). Use 5-10 for initial testing, 10-20 for production runs, 30-50 for comprehensive sampling of challenging targets.
- Timesteps: Number of diffusion denoising steps from random noise to final structure. Default: 50 (optimal quality-speed tradeoff). Range: 20-200 (20 steps provides 10x speedup with equivalent quality; 100-200 shows diminishing returns). Increase for complex topology requirements or final publication-quality runs.
- Hotspots (binder mode): Interface residues on target protein that binder should contact. Format: comma-separated single residues (A50,A51,A52). Ranges like A50-64 are not supported; expand them explicitly (A50,A51,...,A64). Biases diffusion to create contacts with specified residues. Use experimentally validated binding sites or predicted epitopes; leaving blank generates binders to any surface region.
- Binder length (binder mode): Size range for designed binder protein. Range: 5-200 residues (peptides to small domains). Typical values: 40-80 residues for stable single-domain binders. Smaller binders are easier to produce but may have lower affinity.
- Partial diffusion timesteps: Controls diversity when partially redesigning structures. Range: 0-50 (zero = no change, 50 = complete redesign). Auto mode automatically determines noising level. Manual values: 10-20 for subtle variations, 25-30 for moderate diversity (most common), 40-50 for radical changes.
- Total assembly length (unconditional mode): Total size of the de novo generated protein or symmetric assembly. Range: 10-500 residues. For symmetric assemblies, the value must divide evenly across the generated subunits. Practical range: 50-200 residues for well-folded single domains (larger proteins may have lower experimental success rates).
- Symmetry (unconditional mode): Generates symmetric oligomeric assemblies. Options: none (monomer), cyclic (Cn), dihedral (Dn), tetrahedral (T), octahedral (O), icosahedral (I). Symmetry order controls cyclic Cn and dihedral Dn notation; dihedral assemblies generate 2n subunits. Applications include protein cages, virus-like particles, and multivalent binders.
- Binding pocket (binder mode): Crops target structure to focus on specific region. Format: residue range (e.g., 50-150). Reduces computational complexity and focuses design on relevant surface. Use for large target proteins where binding site is localized.
- Motif chain (scaffolding mode): Specifies which chain contains functional motif to scaffold. Format: single chain ID (A, B, C, etc.). Identifies residues to preserve exactly. Common with enzyme active sites, binding epitopes, and metal coordination sites.
- Scaffold extensions (scaffolding mode): Defines how much new structure to add around motif. N-terminal and C-terminal ranges specify residues to add (e.g., 5-15 and 10-20). Minimum 5 residues for structural stability; 10-20 residues typical for well-folded scaffolds.
- Contig string (custom mode): Domain-specific language for precise control over structure generation. Syntax:
A10-100/0 50-150keeps A10-100, breaks chain, generates 50-150 residue chain. Advanced users requiring exact topological specifications. See RFdiffusion GitHub for full syntax. - Use beta model: Binder-design checkpoint with improved secondary structure element balance. Default: off (RFdiffusion selects its normal checkpoint from the submitted config). Enable for binder jobs when outputs show excessive alpha-helices or need better balance of sheets and helices.
- Cyclic chains: Creates macrocyclic structures with covalent N-C termini connection. Default: off (linear chains). Applications include improved stability, constrained conformations, and therapeutic peptides. Sequence design must accommodate cyclization.
- Guiding potentials (experimental): Biases diffusion toward desired biophysical properties—start with defaults before experimenting. Options: monomer ROG (compact structures), monomer contacts (intra-chain stability), oligomer contacts (multi-subunit interfaces), substrate contacts (binding sites around ligands), binder-specific potentials. Warning: can degrade quality if misused; mode-specific dependencies apply.
Understanding the results
Native output files
RFdiffusion returns its native backbone design files:
- Design PDB files: Final RFdiffusion backbone predictions. Designed residues are output as glycine because RFdiffusion performs backbone generation, not sequence design or side-chain packing.
- TRB metadata files: Native run metadata for each design, including the sampled contig, RFdiffusion config, pLDDT trajectory values, device, runtime, and residue mapping arrays.
- Trajectory PDB files: Native
Xt-1andpX0multi-step trajectory files when trajectory output is enabled by RFdiffusion. These show the denoising path and can be inspected in molecular viewers that support multi-model PDBs.
Run ProteinMPNN, LigandMPNN, AlphaFold, or a related validation workflow after RFdiffusion when you need sequence design, side-chain packing, or fold-confidence metrics.
Use cases
Protein binder design
High-affinity binders to therapeutic targets represent the most experimentally validated application. Published examples include nanomolar binders to MDM2 (0.5-0.7 nM vs 600 nM native), influenza hemagglutinin, IL-7Rα, PD-L1, and TrkA. Independent validation by Adaptyv Bio on IL-7Rα demonstrated 84% confirmation rate with 23/27 binders showing measurable affinity (strongest: 40 nM).
Cryo-EM structures of designed binders show near-perfect agreement with computational models. Some designed interfaces achieve picomolar affinity through pure computation without experimental optimization.
Enzyme active site scaffolding
RFdiffusion can scaffold catalytic residues into novel protein folds with specified symmetry. The method enables de novo enzyme design by building NTF2-like folds around functional sites. Applications include designing protein-metal assemblies around coordination sites and transferring catalytic motifs into alternative structural contexts.
Symmetric oligomer design
Generates novel symmetric assemblies validated by electron microscopy across all symmetry types—cyclic (Cn), dihedral (Dn), tetrahedral (T), octahedral (O), and icosahedral (I). Examples include C3 symmetric trimers targeting SARS-CoV-2 spike protein. Hundreds of designed metal-binding symmetric proteins have been experimentally characterized.
De novo protein generation
Unconditional mode creates entirely novel protein folds not seen in nature. No template or evolutionary information required. Topology-constrained monomer designs demonstrate the model's ability to explore uncharted regions of protein structure space.
Success rates
Original Nature paper: 55/96 designs (57%) showed detectable binding at 10 μM, representing ~2 orders of magnitude improvement over previous methods. However, recent critical evaluations show variable success for challenging eukaryotic targets, highlighting the need for generating hundreds to thousands of designs for difficult cases.
Newer methods like BindCraft and EvoPro achieve ~50% success rates (>10x improvement over earlier approaches), while specialized tools show varying performance: RFpeptides (macrocycles) 1.72%, Latent-X 8.26%.
Best practices
Getting started
Start with unconditional mode (100-150 residues) to understand outputs and quality metrics. Use default parameters initially: 10 designs, 50 timesteps, standard temperature settings. Begin with simple design tasks before adding guiding potentials or complex contigs.
Binder design workflow
- Prepare target structure: Clean PDB, remove waters, ensure proper protonation
- Identify binding site: Use hotspots if experimentally known
- Generate 10-20 initial backbone designs with default parameters
- Run sequence design and fold prediction before filtering by pAE_interaction, pLDDT, ipTM, or related downstream metrics
- Validate top 3-5 candidates with MD simulations
- Plan experimental testing: Expect 10-100x designs needed vs confirmed hits
Motif scaffolding strategy
Define motif precisely with exact residue ranges. Allow sufficient N/C-terminal extensions (minimum 5-15 residues) for structural stability. Use substrate_contacts guiding potential if ligand is present. Validate that motif geometry is preserved in outputs by measuring RMSD.
Partial diffusion approach
Start with partial_T=25-30 for moderate diversity around existing fold. Increase to 40-50 only when more radical structural changes are needed. Use provide_seq parameter to retain sequences of critical functional residues. Check RMSD to input structure to ensure appropriate diversity level.
Parameter tuning guidance
Increase timesteps when facing complex topology requirements, poor initial results at 50 steps, or preparing final production runs. Enable beta model when outputs show excessive helices or need better secondary structure balance. Add guiding potentials only after trying without them first—use for specific biophysical requirements informed by experimental feedback.
Validation pipeline
- Sequence and fold prediction: Run ProteinMPNN or LigandMPNN, then evaluate AlphaFold pLDDT, pAE, ipTM, and pTM metrics
- Structural: Run MD simulations, Rosetta relaxation
- Biophysical: Compare AlphaFold2 multimer predictions across candidate sequences
- Experimental: Expression testing, purification, binding assays
- High-resolution: X-ray crystallography or cryo-EM for final candidates
Common pitfalls to avoid
Skipping downstream fold-confidence checks can hide weak designs. Testing only the rank 1 design misses diversity—examine top 3-5. Insufficient sampling (too few designs) reduces statistical power for finding successful candidates. Skipping proper protein preparation leads to artifacts. Over-reliance on any single metric misses important failure modes captured by complementary scores.
Limitations
Technical constraints
Standard RFdiffusion performs backbone-only design without explicit side chain or ligand modeling (use LigandMPNN for post-processing). The target protein is treated as rigid during design. Water molecules and explicit solvation are not modeled. Computational cost requires a 20-minute inference timeout and 250 credits per job on ProteinIQ.
Biological challenges
Success rates vary from 1-50% depending on target difficulty. Designed sequences may show low recombinant expression in standard systems. Some designs exhibit non-specific binding or promiscuity. Affinity ranges cluster in hundreds of nanomolar, often requiring optimization for therapeutic applications.
Applicability limits
Primarily designs with canonical amino acids. Practical size constraints: 10-500 residues. Novel fold generation remains challenging despite improvements over alternatives. Metalloproteins require special consideration. Covalent modifications are not directly supported.
Model uncertainty
Confidence scores are predictions, not guarantees of experimental success. Training on x-ray crystal structures may reduce performance on computational models. Domain generalization struggles with highly novel protein families. Limited training data for macrocycles and non-standard protein architectures.
Cost
ProteinIQ pricing: 250 credits base cost per job. 20-minute inference timeout limit. Moderate-to-high cost tool reflecting complex AI model and GPU-intensive computation.
Optimization: Start with 10 designs (not 50), use 50 timesteps (not 200), test small-scale before large campaigns, batch related designs together.
Related tools

EvoDiff
EvoDiff is a diffusion-based protein sequence generation framework from Microsoft Research. ProteinIQ currently runs the EvoDiff-Seq OA_DM_38M model for unconditional protein generation, motif scaffolding, and user-sequence inpainting.

Genie 3
Generate protein structures and scaffolds with Genie 3, an all-atom SE(3)-equivariant diffusion model. Genie 3 supports unconditional protein generation, motif scaffolding, and hotspot-targeted binder design.

ODesign
All-atom generative AI for designing protein binders. Specify target binding sites and generate diverse binding proteins with fine-grained control over interaction parameters.

PocketFlow
PocketFlow is a structure-based molecular generative model that designs novel drug-like molecules within protein binding pockets. It uses autoregressive flow modeling with chemical knowledge to generate 100% chemically valid, highly drug-like compounds.

PocketXMol
PocketXMol is a pocket-interacting generative foundation model for small-molecule or peptide docking and design in protein binding pockets.

Proteo-R1
Exploratory antibody CDR co-design for antibody-antigen complexes using Proteo-R1 reasoning and raw diffusion. The standard online workflow does not include the framework structure-inpainting assets required for the published-quality target.

RFdiffusion 2
RFdiffusion2 is an atom-level enzyme active site scaffolding tool that generates protein scaffolds around your input motif. REQUIRES an input PDB structure containing the active site residues to scaffold. For ligand-aware design, ligands must be embedded in the input PDB as HETATM records.

ProGen2
ProGen2 is Salesforce Research's protein language model suite for prompt-based de novo protein sequence generation and bidirectional sequence likelihood scoring.

BoltzGen
BoltzGen uses generative diffusion models to design protein, peptide, nanobody, and Fab binders against protein and small-molecule targets.

PepMimic
PepMimic designs short peptides that mimic the binding interface of a known protein binder on its target. From a reference protein complex, a latent diffusion model generates peptide candidates constrained to the target interface, and each candidate is scored by interface-mimicry against the reference binder.