Use case
Virtual screening
Compare in silico screening approaches, then start an online workflow for compound preparation, ranking, review, and export.
Structure-based virtual screening
Uses a three-dimensional target to generate and score candidate binding poses, often alongside pocket, property, and pose-quality review.
Shape-based virtual screening
Ranks candidate conformers by three-dimensional overlap with one or more reference ligands, optionally including chemical-feature similarity.
Ligand-based virtual screening
Ranks compounds using similarity, molecular fingerprints, pharmacophores, or learned features derived from known ligands.
Pharmacophore-based virtual screening
Searches for compounds that match a three-dimensional arrangement of interaction features such as donors, acceptors, aromatic regions, hydrophobic regions, and charge centers.
High-throughput virtual screening
Applies staged filtering, batching, and compute planning to screen much larger libraries while keeping throughput and failure handling explicit.
Inverse virtual screening
Evaluates one compound against many potential targets to generate hypotheses about intended targets, off-targets, selectivity, or repurposing opportunities.
What is virtual screening?
Virtual screening is the computational prioritization of compounds before physical testing. A campaign applies one or more ranking methods to a chemical library, then selects a smaller set for review or experiments. The output is a method-dependent shortlist, not a confirmed set of binders, active compounds, or safe candidates.
The starting evidence determines the method. Structure-based screening evaluates candidates in a target structure, while ligand-based, shape, and pharmacophore methods start from known ligands or interaction features. High-throughput screening describes how those calculations are staged and scaled; inverse screening instead starts from one compound and compares potential targets.
Useful campaigns define preparation rules, controls, score direction, failure handling, and validation before the full library runs. Sequential funnels can reserve slower methods for a smaller subset, while parallel screens preserve independent rankings. Either design should retain the structures, settings, exclusions, and evidence behind every shortlisted compound.
When to use virtual screening
- Prioritize a compound library. Reduce a broader library to a smaller, reviewable set before committing compounds to more expensive calculations or assays.
- Review candidates against a defined target. Generate poses and compare target-specific docking evidence when a suitable structure and binding-site context are available.
- Build an inspectable evidence trail. Keep preparation, scoring, property review, pose checks, predictive endpoints, and exported files connected for handoff and follow-up.
Benefits of virtual screening
- Narrows experimental choices. A defined computational funnel can reduce a broad library to a smaller set that is practical to purchase, synthesize, and test.
- Uses the evidence already available. Programs can start from a target structure, known ligands, interaction features, or a combination without forcing every project into docking.
- Makes prioritization inspectable. Retained structures, scores, alignments, poses, failures, settings, and exclusions show why each candidate advanced.
- Supports staged validation. Fast early methods can reserve slower calculations and expert review for a focused subset before experiments.
Primary limitations
- Every ranking inherits model bias. Reference ligands, target structures, decoys, training data, and scoring functions define what the screen can recognize.
- Preparation choices change results. Salts, stereochemistry, protonation, tautomers, conformers, receptor state, and pocket definition can alter ranks substantially.
- Different scores are not interchangeable. Similarity, docking, pharmacophore fit, and predictive endpoints require their own scales, controls, and interpretation.
- A shortlist is not experimental evidence. Computational prioritization does not confirm binding, activity, selectivity, exposure, safety, or efficacy.
Types of virtual screening
The main types of virtual screening differ in the evidence they start from, the scale they are designed for, and how their rankings should be interpreted.
Structure-based virtual screening
Uses a three-dimensional target to generate and score candidate binding poses, often alongside pocket, property, and pose-quality review.
Best for: A defined target with a suitable structure and a binding site that can support a consistent docking setup.
Requires: A prepared target structure, a compound library, and explicit docking and review settings.
Shape-based virtual screening
Ranks candidate conformers by three-dimensional overlap with one or more reference ligands, optionally including chemical-feature similarity.
Best for: Programs with a bioactive reference ligand or credible bound conformation but no target structure suitable for docking.
Requires: A reviewed reference conformation, a conformer-generation strategy, and a standardized candidate library.
Ligand-based virtual screening
Ranks compounds using similarity, molecular fingerprints, pharmacophores, or learned features derived from known ligands.
Best for: Programs with known active ligands but no sufficiently reliable target structure for structure-based screening.
Requires: One or more relevant reference ligands or an activity set, plus a standardized compound library.
Pharmacophore-based virtual screening
Searches for compounds that match a three-dimensional arrangement of interaction features such as donors, acceptors, aromatic regions, hydrophobic regions, and charge centers.
Best for: Programs with a supported ligand- or structure-derived feature hypothesis that should retrieve chemotypes beyond one molecular scaffold.
Requires: A reviewed pharmacophore model, candidate conformers, explicit feature tolerances, and a validation set or other basis for choosing the model.
High-throughput virtual screening
Applies staged filtering, batching, and compute planning to screen much larger libraries while keeping throughput and failure handling explicit.
Best for: Broad early triage where library size requires a deliberately scaled screening and review strategy.
Requires: A standardized large library, defined throughput limits, reproducible batching, and a plan for reviewing ranked results.
Inverse virtual screening
Evaluates one compound against many potential targets to generate hypotheses about intended targets, off-targets, selectivity, or repurposing opportunities.
Best for: Target identification, off-target investigation, selectivity profiling, or mechanism-of-action hypotheses for a defined compound.
Requires: A reviewed query compound, a curated target panel, comparable target preparation, and a plan for handling score comparability and experimental follow-up.
Virtual screening in drug discovery
Virtual screening in drug discovery is usually an early prioritization step between assembling a chemical library and selecting compounds for physical testing. It can reduce a large search space to a reviewable shortlist, but the value of that shortlist depends on the biological question, input evidence, preparation rules, ranking method, and validation design.
In silico virtual screening may start from a target structure, known ligands, a three-dimensional feature model, or a scale-driven combination of fast and slow methods. An online virtual screening workflow helps keep those inputs, settings, failures, rankings, and exported files connected; it does not make the computational result equivalent to an assay.
Sequential, parallel, and hybrid screening strategies
A screening strategy describes how evidence is combined, not a new score. Choose the design according to the cost of each method, the evidence available at the start, and whether the program needs one filtered ranking or several independent views.
- Sequential funnel. Apply inexpensive filters first, then send progressively smaller subsets to conformer generation, docking, rescoring, or expert pose review.
- Parallel comparison. Run complementary methods independently and retain each native ranking so compounds supported by different evidence are not discarded prematurely.
- Hybrid model. Combine ligand and target information inside one validated method only when its training domain, score meaning, and failure behavior are understood.
How to do virtual screening online
An online run should begin with a decision and validation plan, not with the largest available library. The following sequence keeps the scientific boundary and the operational record visible.
- Choose the screening question. Define the target, reference evidence, library scope, desired shortlist size, and experiment the ranking will inform.
- Prepare a validation subset. Assemble relevant known actives, inactives or decoys when available, plus representative library chemistry and difficult input cases.
- Standardize the inputs. Preserve identifiers while documenting salts, stereochemistry, protonation, tautomers, receptor state, and rejected records.
- Run the matched workflow. Select the structure-, ligand-, shape-, pharmacophore-, high-throughput-, or inverse-screening path and keep method-specific settings fixed for comparison.
- Inspect rankings and failures. Review score distributions, poses or alignments, controls, missing outputs, property context, and chemical diversity before applying a cutoff.
- Export and validate the shortlist. Retain the full run record, then test selected compounds with an orthogonal calculation and experiments appropriate to the biological question.
How a virtual screening workflow works
The featured workflow runs a learned target-aware ranking and structure-based docking as separate evidence branches. Use a method-specific workflow instead when the program lacks a suitable target structure or protein sequence.
- Provide target evidence. Submit a compatible protein sequence for SPRINT and a reviewed PDB structure for the structure-based branch.
- Run target-aware ranking. Rank the submitted SMILES library with SPRINT and retain its native cosine similarity alongside compound identifiers.
- Prepare and dock. Repair the receptor with PDBFixer and dock the library independently with GNINA.
- Review independent evidence. Check GNINA geometry with PoseBusters, calculate residue-level interactions with ProLIF, and review descriptors and ADMET predictions.
- Compare and shortlist. Compare the branches without treating their raw scores as interchangeable, then export a diverse shortlist for orthogonal review.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Target structure.
PDBThe main template accepts a PDB receptor. If the source structure is CIF, convert it to PDB before starting the run. - Compound library.
SMILESSubmit a SMILES library directly. Convert existing SDF with SDF to SMILES or MOL2 with MOL2 to SMILES before launching; use the ligand-preparation variant to generate or repair structure files from SMILES. - Binding-site context. Known pocket coordinates or a reference ligand can inform manual docking setup and review, but they are not required inputs on this template.
Outputs
- Prepared target and pocket results.
PDBJSONDownload the fixed receptor and preparation provenance, plus fpocket pocket files and metrics. - Property and rule tables.
CSVJSONSMILESInspect per-compound Lipinski and lead-likeness results with the values and warnings returned by each step. - Docking poses and scores.
SDFCSVJSONReview ranked docking poses with GNINA CNN score, CNN affinity (pKd), and Vina score (kcal/mol). - Pose and ADMET review.
CSVJSONInspect downloadable PoseBusters validation tables and ADMET-AI prediction tables as separate evidence layers. - Run record.
LOGFILESRetain settings, node-level results, logs, and downloadable files for review or handoff.
Featured virtual screening workflow
SPRINT cosine similarity and GNINA docking outputs remain separate because they have different meanings. PoseBusters and ProLIF review GNINA poses, while molecular descriptors and ADMET-AI provide independent compound-level context.
Inputs
3 required
Methods
7 connected
- 01SPRINT Ranking
- 02PDBFixer
- 03Molecular descriptors
- 04GNINA Docking
- 05ADMET-AI
- 06PoseBusters
- 07ProLIF
SPRINT cosine similarity and GNINA docking outputs remain separate because they have different meanings. PoseBusters and ProLIF review GNINA poses, while molecular descriptors and ADMET-AI provide independent compound-level context.
Use this templateTools for virtual screening
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

ChEMBL Download
Retrieves compounds and activity records for reference sets and screening libraries.

PubChem Download
Downloads compound records and structures from PubChem for library assembly.

Molecular descriptors
Calculates physicochemical descriptors, Morgan fingerprints, and MACCS keys.
Open Babel
Converts, prepares, and standardizes molecular files across common formats.

SMILES to SDF
Generates SDF structures from SMILES for structure-based screening stages.

Veber's rule
Reviews molecular flexibility and polar surface area against Veber criteria.

PAINS filter
Flags PAINS substructures for compound-library review.

AutoDock-GPU
Runs GPU-accelerated docking for higher-throughput structure-based screens.

SMINA
Supports docking, minimization, and alternative scoring for comparative review.

GNINA
Generates ranked docking poses with CNN and Vina scoring outputs.

Admetica
Predicts ADMET properties for post-screening prioritization.

ToxPred 2.0 (Toxicity prediction)
Predicts toxicity risk and reports structural alerts for shortlisted compounds.
Frequently asked questions
Start with a small, representative set that includes compounds with known or expected behavior when possible. Review target and compound preparation, binding-site context, failure rates, pose quality, and ranking behavior before committing a larger library to the same setup.
Current public examples range from $500 for a limited university-core computational screen to $5,000 for a published 1.76-million-compound 3D shape-screening package, $15,000 for an integrated target-assessment project that includes virtual screening, and $20,000 for one service exploring up to 10 billion compounds.
These are unlike deliverables rather than a universal per-compound rate. Library preparation, method, target or reference complexity, compute settings, shortlist analysis, compound delivery, and experimental validation account for much of the difference.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with the configured workflow quoted in credits before it runs. Done-for-you projects are scoped separately.
Yes. Workflow templates are editable, so you can replace compatible preparation, docking, filtering, or prediction steps and adjust their settings. Keep engine-specific scores and assumptions visible when comparing alternative workflow versions.
They can support a consensus strategy, but raw scores from unrelated methods should not be treated as directly interchangeable. Define the combination rule in advance, preserve each method’s original score, and check whether the resulting consensus improves ranking on an appropriate validation set.
Use docking scores for prioritization within a defined setup, not as measured binding affinities or binding free energies. Compare engines in their own scoring context. Treat ADMET and toxicity predictions as triage signals with model-specific uncertainty, not proof of safety or clinical behavior.
Review the underlying structures, poses, model applicability, and failure checks before selecting compounds for follow-up. Shortlisted compounds still require orthogonal computational review and experimental assays appropriate to the target, mechanism, and development question.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.