Protenix v2 icon

Protenix v2

(2.0.0)

Enhanced open-source biomolecular structure prediction with Protenix v2 Learn more

Input

Add a molecule to begin

Choose a building block to assemble your structure.

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job” — or start from an example.

Barnase–barstar protein complex

HIV-1 protease with darunavir

MS2 coat protein bound to RNA

What is Protenix v2?

Protenix v2 is ByteDance Research's enhanced open-source biomolecular structure prediction model, released under the Apache 2.0 license as part of the Protenix project. It predicts 3D structures of complexes containing proteins, DNA, RNA, small molecule ligands, and metal ions from sequence or structural inputs, all in a single run.

AlphaFold 3 itself has restricted access and prohibits commercial use. Protenix follows the same general all-atom structure-prediction direction with open-source code and Apache-licensed model parameters, making it a practical option for open-ended research and applications where AlphaFold 3's terms would be limiting. ProteinIQ now runs Protenix v2 by default while keeping v1, mini, and tiny checkpoints available for comparison or faster exploratory runs.

How to use Protenix v2 online

ProteinIQ runs Protenix v2 on GPU infrastructure with no local installation. Add one or more molecules, optionally configure seeds and model, then submit. Results arrive as downloadable CIF files ranked by confidence score, viewable directly in the 3D structure viewer. Each ranked prediction also includes the Protenix summary-confidence JSON, and you can optionally export atom-level confidence JSON for deeper inspection. MSA search is off by default; enabling it can add several minutes but may improve protein predictions that benefit from evolutionary context.

Inputs

Each molecule is added as a separate entity. Up to 10 entries of each type are supported.

InputAccepted formats
ProteinFASTA sequence, PDB file, or 4-letter RCSB PDB ID
LigandSMILES string, CCD code (e.g. ATP), SDF/MOL file, or PubChem compound ID
DNAFASTA sequence (A, T, G, C)
RNAFASTA sequence (A, U, G, C)
IonCCD code (e.g. ZN, MG, CA, FE)
Protein paired MSAOptional uploaded or pasted .a3m alignment attached to a protein chain label
Protein unpaired MSAOptional uploaded or pasted .a3m alignment attached to a protein chain label
RNA MSAOptional uploaded or pasted .a3m alignment attached to an RNA chain label

For ligand files, Protenix requires a 3D conformer. 2D SDF files (flat structure with no z-coordinates) will fail at featurization.

Precomputed alignments attach to the chain letters shown on the main protein or RNA molecule cards. If there is only one eligible protein or RNA chain, you can leave the MSA target label blank and ProteinIQ will attach it automatically. Protein paired MSA and protein unpaired MSA are separate native inputs; RNA only supports unpaired RNA-MSA in this pass. Template files are still out of scope here.

Settings

SettingDescription
Model seedsNumber of random seeds (1-5, default 1). Each seed independently samples the diffusion trajectory and produces a distinct set of structures.
ModelModel variant. Protenix v2 is the default. See model options below.
Use MSAWhether to run multiple sequence alignment for protein chains (default off). Adds 2-5 minutes but can improve accuracy, especially for proteins with many homologs.
Samples per seedStructures generated per seed (1-5, default 5). Higher values increase output diversity.
PrecisionBF16 (default, faster) or FP32 (full precision). Negligible accuracy difference on A10G hardware.
Training-Free GuidanceApplies physical constraints (steric, torsion, distance) during diffusion sampling to improve geometric plausibility. Off by default; adds compute overhead.
Export atom confidence JSONSaves the full_data JSON for each prediction. Useful for debugging or downstream analysis, but larger than the default summary metrics.

Model options

ModelParametersTraining cutoffWhen to use
Protenix v2464M2021-09-30Default. Highest-accuracy model for most targets, with stronger antibody-antigen and ligand plausibility performance reported by ByteDance.
Protenix v1 base368M2021-09-30Baseline v1 model for comparison and lower compute cost.
Protenix v1 base (2025 data)368M2025-06-30Targets with templates deposited after 2021.
Protenix mini135M2021-09-30Rapid screening, large sequences, or quick estimates before a full run. Roughly 40% faster.

Protenix v1 vs Protenix v2

Protenix v2 is the default model on ProteinIQ. It is a 464M-parameter model released in April 2026, with the same 2021-09-30 training cutoff used for fair comparison to earlier Protenix models.

The Protenix documentation reports stronger antibody-antigen performance and improved ligand plausibility compared with Protenix v1, while retaining MSA, RNA MSA, and template support. It also works with the Training-Free Guidance features introduced in the Protenix 2.0.0 package.

Protenix v1 remains available as a lower-compute comparison point. Its public checkpoints include the base 368M models, the 2025 cutoff model, and smaller mini/tiny variants.

How does Protenix work?

Protenix implements the AlphaFold 3 architecture, which replaces the iterative refinement of AlphaFold 2 with a diffusion-based structure generation process. Prediction starts from randomly initialized atom coordinates and progressively denoises them into a coherent structure over multiple cycles.

Multiple Sequence Alignment

For protein inputs, Protenix queries external MSA databases to find evolutionarily related sequences. Conserved residues and co-evolving residue pairs reveal spatial constraints that guide the structure prediction. The MSA step connects to an external server and is the primary source of runtime variability, typically adding 2-5 minutes per prediction.

If you already have alignments, ProteinIQ can use precomputed .a3m files instead of forcing a new search. Protein chains can receive separate paired and unpaired MSAs, while RNA chains can receive an unpaired RNA-MSA. These uploads are written to temporary files and referenced through the pairedMsaPath and unpairedMsaPath JSON fields used by Protenix.

Confidence estimation

Protenix outputs per-atom and per-residue confidence estimates derived from the same network that produces the structure. These are not validated against experimental data for each prediction; they reflect the model's internal consistency. A high-confidence prediction that matches an incorrect prior (e.g. a homolog with a different binding mode) may score well while being physically wrong.

Understanding the results

Protenix outputs ranked CIF files together with the native summary-confidence JSON for each prediction.

MetricScopeInterpretation
pLDDTPer-residue, 0-100>90: high confidence. 70-90: backbone likely correct. 50-70: low confidence. <50: likely disordered or unreliable.
pTMWhole structure, 0-1>0.5 suggests the overall fold is correct.
ipTMInter-chain interfaces, 0-1>0.7 indicates a reliable interface prediction. Particularly relevant for protein-protein and protein-ligand complexes.

ProteinIQ also preserves the additional native metrics and flags from summary_confidence, including ranking_score, gpde, per-chain confidence summaries, pairwise interface scores, clash flags, disorder estimates, and the recycle count used for the prediction.

When running multiple seeds, structural agreement across the top-ranked predictions is a better indicator of reliability than any single confidence score. Divergent predictions across seeds suggest the structure is genuinely uncertain.

Protenix v2 examples

These completed jobs show how Protenix v2 handles protein-protein, protein-ligand, and protein-RNA systems. Open any example to inspect all three ranked predictions and download the CIF and confidence files.

Barnase-barstar protein complex

This example predicts the well-characterized complex between the 110-residue barnase ribonuclease and its 89-residue barstar inhibitor. It demonstrates a compact protein-protein interface where evolutionary context materially improves the prediction.

  • Inputs: Barnase chain A and barstar chain B, 199 residues total
  • Non-default settings: Use MSA = on to supply evolutionary context; Samples per seed = 3 to return a concise ranked set; Export atom confidence JSON = on for residue- and atom-level inspection
Protenix v2 Structure view of the barnase-barstar complex with three predictions ranked at 0.975
Protenix v2 Structure view of the barnase-barstar complex with three predictions ranked at 0.975

The top prediction has a ranking_score of 0.975, pLDDT of 96.0, pTM of 0.977, and ipTM of 0.975, with no clash flag. The three displayed ranking scores round to the same value, showing consistent sampling for this run. These scores indicate strong model confidence in the fold and interface, but they do not replace comparison with an experimental complex structure.

HIV-1 protease with darunavir

This example submits the HIV-1 protease sequence as two copies together with stereospecific darunavir SMILES. It is useful for seeing how a protein-ligand prediction can return a plausible protein fold while remaining uncertain about the interaction.

  • Inputs: HIV-1 protease homodimer and darunavir as a stereospecific SMILES string
  • Non-default settings: Samples per seed = 3 to compare alternative predictions; Export atom confidence JSON = on for detailed confidence analysis
Protenix v2 Structure view of HIV-1 protease with darunavir and ranking scores from 0.533 to 0.569
Protenix v2 Structure view of HIV-1 protease with darunavir and ranking scores from 0.533 to 0.569

The top prediction has a ranking_score of 0.569, pLDDT of 80.0, pTM of 0.713, and ipTM of 0.533. The orange ligand is visible near the protein surface, but the interface score is below the 0.7 confidence heuristic used above. Treat this as an exploratory binding hypothesis: the run does not establish the experimental pose, binding affinity, or antiviral activity.

MS2 coat protein bound to operator RNA

This protein-RNA example combines two copies of the MS2 bacteriophage coat protein with its 19-nucleotide operator hairpin. It demonstrates joint prediction across polymer types without an MSA search.

  • Inputs: MS2 coat-protein homodimer and the ACAUGAGGAUCACCCAUGU operator RNA
  • Non-default settings: Samples per seed = 3 to compare the sampled complexes; Export atom confidence JSON = on for detailed confidence analysis
Protenix v2 Structure view of the MS2 coat protein-RNA complex with a top ranking score of 0.809
Protenix v2 Structure view of the MS2 coat protein-RNA complex with a top ranking score of 0.809

The RNA hairpin extends across the predicted protein complex in the Structure view. The top result has a ranking_score of 0.809, pLDDT of 81.2, pTM of 0.797, and ipTM of 0.812; the other two ranking scores are 0.804 and 0.802. The high interface score and agreement across the three samples support this model as a useful structural hypothesis, not experimental confirmation of the RNA-binding geometry.

When to use Protenix v2 vs alternatives

The main differences are molecule scope, evolutionary inputs, guidance options, and whether the result describes one structure, an ensemble, or binding affinity.

ToolBest forKey distinction
Protenix v2General all-atom complexesMultiple model sizes, protein and RNA MSAs, optional atom confidence; no affinity prediction or user restraints.
Boltz-2Protein-ligand structure and affinityReturns binder probability and affinity estimates; supports templates, constraints, and CCD covalent bonds.
Chai-1Guided complex predictionCombines MSA and template search with contact, pocket, and modified-residue inputs.
OpenFold 3Open AlphaFold 3 research workflowsTrainable research preview with automatic MSA generation and low-memory inference.
RosettaFold3Alternative all-atom complex predictionHandles proteins, DNA, RNA, and ligands with multi-model sampling.
IntelliFold 2Fast multi-component predictionOffers a fast Flash model, optional MSA generation, and ranked complex structures.
AlphaFold2Protein monomers and multimersEstablished protein-only workflow with single-sequence or MSA-assisted prediction.
ESMFold2Protein chains and protein complexesLanguage-model folding with no MSA search and optional PAE and pair-chain ipTM outputs.
ESMFoldRapid single-sequence protein foldingPredicts protein structures without an MSA and returns PDB, pLDDT, and optional PAE.
AlphaFlowProtein conformational ensemblesSamples multiple structures to represent flexibility rather than one static prediction.
HighFoldCyclic peptide structuresModels head-to-tail cyclization and disulfide constraints with an AlphaFold2-derived method.

Protenix v2 is a strong default for unconstrained complexes containing several molecule types. Boltz-2 is the clearer choice for ligand affinity or pocket guidance, while Chai-1 is useful when templates, known contacts, or modified residues should influence the model.

Folding models predict structures from sequences and chemical inputs. For docking into a prepared receptor, AutoDock Vina searches a defined box and reports scores in kcal/mol, while DiffDock predicts ligand poses without a predefined box.

Table of contents

Related tools

Chai-1

Chai-1

Chai-1 is a multi-modal foundation model for molecular structure prediction. Predicts 3D structures for proteins, ligands, DNA, RNA, and multi-component complexes with high accuracy.

protein-foldingstructure-prediction+5
IntelliFold 2

IntelliFold 2

Controllable all-atom structure prediction for proteins, ligands, DNA, RNA, and multi-component complexes using IntelliFold 2.0.4 on its AlphaFold 3 JAX engine.

protein-foldingai-powered+5
OpenFold-3

OpenFold-3

OpenFold-3 is an open-source AI model for biomolecular structure prediction, aiming to reproduce AlphaFold3. Predicts 3D structures for proteins, RNA, DNA, and small molecule ligands with high accuracy.

protein-foldingstructure-prediction+5
RosettaFold3

RosettaFold3

Open-source structure prediction neural network for proteins, nucleic acids, and small molecules. State-of-the-art accuracy with multi-chain support.

protein-foldingstructure-prediction+5
Boltz-2

Boltz-2

Boltz-2 is a biomolecular foundation model for structure and binding affinity prediction. Supports proteins, ligands, DNA, and RNA in multi-component complexes. Automatically scales GPU resources for large complexes. Predicts binding affinity with near-FEP accuracy at 1000x faster speed.

protein-foldingstructure-prediction+5
LMI4Boltz

LMI4Boltz

LMI4Boltz is a low-memory fork of Boltz for biomolecular structure and binding affinity prediction. It preserves Boltz inference behavior while reducing VRAM use with in-place pair updates, CPU offload, reduced precision pair representation, and aggressive chunking.

protein-foldingstructure-prediction+5
ABodyBuilder3

ABodyBuilder3

ABodyBuilder3 predicts antibody variable-domain structures from paired heavy and light chain sequences. It returns a PDB structure and, for the pLDDT checkpoint, per-residue confidence values.

protein-foldingstructure-prediction+3
AlphaFlow

AlphaFlow

Generate protein conformational ensembles with AlphaFlow or ESMFlow from a sequence, optional MSA, and supported reference-structure checkpoints.

protein-foldingstructure-prediction+4
AlphaFold2

AlphaFold2

AlphaFold2 via ColabFold for protein structure prediction. Free runs use single-sequence mode; paid plans add MMseqs2 MSA generation. Supports monomer and multimer prediction.

protein-foldingstructure-prediction+4
ESMfold

ESMfold

ESMfold is a fast, single-sequence protein structure predictor from Meta AI. Predicts 3D protein structures directly from amino acid sequences without requiring multiple sequence alignments (MSA), making it significantly faster than AlphaFold while automatically scaling GPU resources for larger proteins.

protein-foldingstructure-prediction+2