Use case
Consensus docking
Compare results across docking engines while preserving each method’s native poses, scores, and confidence values.
Inputs
2 required
Methods
3 connected
- 01AutoDock Vina
- 02GNINA
- 03DiffDock-L
Run AutoDock Vina, GNINA, and DiffDock-L in parallel while preserving method-native results for explicit comparison.
Use this templateWhat is consensus docking?
Consensus docking is a method for combining or comparing results from multiple docking engines or scoring functions. It identifies predictions that remain stable across methods, but a defensible protocol defines pose matching and rank aggregation in advance rather than averaging incompatible raw scores.
Consensus docking compares multiple docking engines, search strategies, or scoring models run on the same prepared inputs. Agreement across independent methods may reduce sensitivity to one engine’s assumptions. Pose agreement, rank aggregation, interaction-pattern comparison, and post-docking rescoring are separate analyses and need separate rules.
Do not average AutoDock Vina energies, neural-network scores, and diffusion confidence values as if they measured one quantity. Define how equivalent poses are matched, how rankings are normalized, and how failed results are handled. Several related scoring functions can agree because they share the same bias.
When to use consensus docking
- Suitable consensus docking question. Checking whether a pose or prioritization survives changes in docking method
- Required structures and evidence are available. Identically prepared inputs, multiple complementary engines, and a predefined consensus rule
Benefits of consensus docking
- Practical output. Exposes method sensitivity
- Comparative evidence. Can prioritize recurring poses
- Connected analysis. Reduces reliance on one scoring function
Primary limitations
- Method dependence. Agreement can reflect shared bias
- Input sensitivity. Raw scores are not directly commensurate
- Validation boundary. More engines do not guarantee correctness
How consensus docking works
Consensus protocols should name the object being combined rather than using “consensus” as a generic label.
- Consensus pose selection. Poses from different engines are clustered by ligand geometry or interaction pattern, and recurring binding modes are prioritized for inspection.
- Consensus ranking. Within-method ranks are combined using a declared rule such as rank voting or best-rank selection. This avoids treating incompatible raw scores as one unit.
- Consensus rescoring. A common set of poses is evaluated by several scoring functions. Correlated scores and shared failure modes still need to be considered.
Applications of consensus docking
Consensus docking is useful when method sensitivity is itself an important part of the decision.
- Pose robustness. Identify binding modes that recur across classical, CNN-assisted, and learned pose-generation methods.
- Screening triage. Reduce dependence on one ranking by combining declared within-method ranks, while preserving compounds missing from particular methods.
- Protocol benchmarking. Compare candidate consensus rules on known complexes or actives before applying the rule prospectively.
How to do consensus docking online
ProteinIQ’s consensus workflow runs three complementary engines in parallel and intentionally leaves their native outputs separate for transparent review.
- Standardize one shared input set. Use the same prepared protein, ligand states, cofactors, and search assumptions for every engine so preparation differences do not masquerade as method differences.
- Choose complementary methods. Select engines with meaningfully different search or scoring assumptions. The included workflow compares AutoDock Vina, GNINA, and DiffDock-L.
- Run each method independently. Retain every method-native pose, score, confidence value, log, and failure instead of overwriting results with one merged table.
- Match comparable poses. Cluster poses using ligand heavy-atom geometry, symmetry-aware RMSD, or declared interaction criteria. Do not call two poses a consensus merely because their scores are both favorable.
- Apply and report the consensus rule. Use a predefined pose- or rank-level rule, test sensitivity to missing outputs, and export the underlying per-engine evidence with the consensus result.
How to interpret consensus docking results
Agreement increases robustness to method choice only when the methods provide partly independent evidence. If all engines inherit the same receptor error, incorrect ligand state, or training bias, consensus can reinforce the same wrong answer.
Benchmark the complete decision rule against relevant known complexes or screening data instead of evaluating each engine in isolation. Report disagreements and failures because they are often more informative than a forced combined score.
How the consensus docking workflow works
Run AutoDock Vina, GNINA, and DiffDock-L in parallel while preserving method-native results for explicit comparison.
- Standardize shared inputs. Use the same prepared protein, ligand states, cofactors, and search assumptions for every engine so preparation differences do not masquerade as method differences.
- Choose complementary engines. Select engines with meaningfully different search or scoring assumptions. The included workflow compares AutoDock Vina, GNINA, and DiffDock-L.
- Run independent docking. Retain every method-native pose, score, confidence value, log, and failure instead of overwriting results with one merged table.
- Match comparable poses. Cluster poses using ligand heavy-atom geometry, symmetry-aware RMSD, or declared interaction criteria. Do not call two poses a consensus merely because their scores are both favorable.
- Apply a declared consensus rule. Use a predefined pose- or rank-level rule, test sensitivity to missing outputs, and export the underlying per-engine evidence with the consensus result.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Structural inputs.
PDBSDFSMILESOne consistently prepared protein and ligand set shared by all selected docking engines. - Method context. Binding-site evidence, restraints, receptor-state provenance, known ligands, or reference complexes when available.
Outputs
- Docked structures.
PDBPDBQTSDFPer-engine poses and scores, pose clusters, agreement summaries, and retained method provenance. - Review evidence. Method-native rankings, confidence, logs, interaction context, failures, and files for reproducible follow-up.
Tools for consensus docking
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

AutoDock Vina
Provide a classical docking baseline

GNINA
Add CNN-assisted scoring

DiffDock-L
Add diffusion-based pose prediction

SMINA
Add customizable Vina-derived scoring

AutoDock-GPU
Add AutoDock4 search and scoring

PandaDock
Add a physics-based comparison

TEMPL Pipeline
Add template-guided pose evidence

FlowDock
Add flow-based complex prediction

PoseBusters
Filter implausible poses

ProLIF
Compare interaction fingerprints

PDBFixer
Keep receptor preparation consistent

fpocket
Check pocket context
Other molecular docking workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Protein–ligand docking
Predict small-molecule binding poses in a protein target and inspect scoring and interaction evidence.
Protein–protein docking
Predict how two protein partners assemble into a biomolecular complex.
Antibody–antigen docking
Explore antibody recognition orientations with antibody-aware interface evidence.
Peptide–protein docking
Model a flexible peptide partner against a protein receptor.
Blind docking
Search a protein broadly when the relevant binding site is unknown.
Flexible molecular docking
Account for ligand flexibility and selected or learned receptor movement.
Ensemble docking
Dock against multiple receptor conformations instead of one static structure.
AI molecular docking
Use learned diffusion or flow models to predict protein–ligand complexes.
Frequently asked questions
Usually not. Scores from different engines may use different units, scales, calibration, and meanings. Compare matched poses or within-method ranks, or rescore a shared pose set with a declared protocol; preserve every native value instead of presenting an unsupported arithmetic average.
Checking whether a pose or prioritization survives changes in docking method
One consistently prepared protein and ligand set shared by all selected docking engines.
Accuracy depends on target class, input preparation, conformational coverage, method domain, and the evaluation criterion. Benchmark against relevant known complexes and report pose accuracy separately from ranking or affinity claims.
Public core rates may be approximately $55–$169 per labor hour plus compute, and multi-engine protocols multiply run cost. These public rates are service examples rather than universal prices; scope, preparation, number of systems, methods, compute, interpretation, and experimental work change the total.
ProteinIQ Plus is $29 per month and Pro is $99 per month. Compute-heavy runs also consume credits according to the selected tool and workload; a separately scoped done-for-you engagement is available when experimental design, data preparation, interpretation, or reporting needs expert support.
Consensus is robustness evidence, not truth. Report the matching and aggregation rule, preserve every engine’s native output, and benchmark the protocol against relevant known complexes.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.