What is virtual screening?

Virtual screening is the computational prioritization of compounds before physical testing. A campaign applies one or more ranking methods to a chemical library, then selects a smaller set for review or experiments. The output is a method-dependent shortlist, not a confirmed set of binders, active compounds, or safe candidates.

The starting evidence determines the method. Structure-based screening evaluates candidates in a target structure, while ligand-based, shape, and pharmacophore methods start from known ligands or interaction features. High-throughput screening describes how those calculations are staged and scaled; inverse screening instead starts from one compound and compares potential targets.

Useful campaigns define preparation rules, controls, score direction, failure handling, and validation before the full library runs. Sequential funnels can reserve slower methods for a smaller subset, while parallel screens preserve independent rankings. Either design should retain the structures, settings, exclusions, and evidence behind every shortlisted compound.

When to use virtual screening

  • Prioritize a compound library. Reduce a broader library to a smaller, reviewable set before committing compounds to more expensive calculations or assays.
  • Review candidates against a defined target. Generate poses and compare target-specific docking evidence when a suitable structure and binding-site context are available.
  • Build an inspectable evidence trail. Keep preparation, scoring, property review, pose checks, predictive endpoints, and exported files connected for handoff and follow-up.

Benefits of virtual screening

  • Narrows experimental choices. A defined computational funnel can reduce a broad library to a smaller set that is practical to purchase, synthesize, and test.
  • Uses the evidence already available. Programs can start from a target structure, known ligands, interaction features, or a combination without forcing every project into docking.
  • Makes prioritization inspectable. Retained structures, scores, alignments, poses, failures, settings, and exclusions show why each candidate advanced.
  • Supports staged validation. Fast early methods can reserve slower calculations and expert review for a focused subset before experiments.

Primary limitations

  • Every ranking inherits model bias. Reference ligands, target structures, decoys, training data, and scoring functions define what the screen can recognize.
  • Preparation choices change results. Salts, stereochemistry, protonation, tautomers, conformers, receptor state, and pocket definition can alter ranks substantially.
  • Different scores are not interchangeable. Similarity, docking, pharmacophore fit, and predictive endpoints require their own scales, controls, and interpretation.
  • A shortlist is not experimental evidence. Computational prioritization does not confirm binding, activity, selectivity, exposure, safety, or efficacy.

Types of virtual screening

The main types of virtual screening differ in the evidence they start from, the scale they are designed for, and how their rankings should be interpreted.

Structure-based virtual screening

Uses a three-dimensional target to generate and score candidate binding poses, often alongside pocket, property, and pose-quality review.

Best for: A defined target with a suitable structure and a binding site that can support a consistent docking setup.
Requires: A prepared target structure, a compound library, and explicit docking and review settings.

Shape-based virtual screening

Ranks candidate conformers by three-dimensional overlap with one or more reference ligands, optionally including chemical-feature similarity.

Best for: Programs with a bioactive reference ligand or credible bound conformation but no target structure suitable for docking.
Requires: A reviewed reference conformation, a conformer-generation strategy, and a standardized candidate library.

Ligand-based virtual screening

Ranks compounds using similarity, molecular fingerprints, pharmacophores, or learned features derived from known ligands.

Best for: Programs with known active ligands but no sufficiently reliable target structure for structure-based screening.
Requires: One or more relevant reference ligands or an activity set, plus a standardized compound library.

Pharmacophore-based virtual screening

Searches for compounds that match a three-dimensional arrangement of interaction features such as donors, acceptors, aromatic regions, hydrophobic regions, and charge centers.

Best for: Programs with a supported ligand- or structure-derived feature hypothesis that should retrieve chemotypes beyond one molecular scaffold.
Requires: A reviewed pharmacophore model, candidate conformers, explicit feature tolerances, and a validation set or other basis for choosing the model.

High-throughput virtual screening

Applies staged filtering, batching, and compute planning to screen much larger libraries while keeping throughput and failure handling explicit.

Best for: Broad early triage where library size requires a deliberately scaled screening and review strategy.
Requires: A standardized large library, defined throughput limits, reproducible batching, and a plan for reviewing ranked results.

Inverse virtual screening

Evaluates one compound against many potential targets to generate hypotheses about intended targets, off-targets, selectivity, or repurposing opportunities.

Best for: Target identification, off-target investigation, selectivity profiling, or mechanism-of-action hypotheses for a defined compound.
Requires: A reviewed query compound, a curated target panel, comparable target preparation, and a plan for handling score comparability and experimental follow-up.

Virtual screening in drug discovery

Virtual screening in drug discovery is usually an early prioritization step between assembling a chemical library and selecting compounds for physical testing. It can reduce a large search space to a reviewable shortlist, but the value of that shortlist depends on the biological question, input evidence, preparation rules, ranking method, and validation design.

In silico virtual screening may start from a target structure, known ligands, a three-dimensional feature model, or a scale-driven combination of fast and slow methods. An online virtual screening workflow helps keep those inputs, settings, failures, rankings, and exported files connected; it does not make the computational result equivalent to an assay.

Sequential, parallel, and hybrid screening strategies

A screening strategy describes how evidence is combined, not a new score. Choose the design according to the cost of each method, the evidence available at the start, and whether the program needs one filtered ranking or several independent views.

  • Sequential funnel. Apply inexpensive filters first, then send progressively smaller subsets to conformer generation, docking, rescoring, or expert pose review.
  • Parallel comparison. Run complementary methods independently and retain each native ranking so compounds supported by different evidence are not discarded prematurely.
  • Hybrid model. Combine ligand and target information inside one validated method only when its training domain, score meaning, and failure behavior are understood.

How to do virtual screening online

An online run should begin with a decision and validation plan, not with the largest available library. The following sequence keeps the scientific boundary and the operational record visible.

  1. Choose the screening question. Define the target, reference evidence, library scope, desired shortlist size, and experiment the ranking will inform.
  2. Prepare a validation subset. Assemble relevant known actives, inactives or decoys when available, plus representative library chemistry and difficult input cases.
  3. Standardize the inputs. Preserve identifiers while documenting salts, stereochemistry, protonation, tautomers, receptor state, and rejected records.
  4. Run the matched workflow. Select the structure-, ligand-, shape-, pharmacophore-, high-throughput-, or inverse-screening path and keep method-specific settings fixed for comparison.
  5. Inspect rankings and failures. Review score distributions, poses or alignments, controls, missing outputs, property context, and chemical diversity before applying a cutoff.
  6. Export and validate the shortlist. Retain the full run record, then test selected compounds with an orthogonal calculation and experiments appropriate to the biological question.

How a virtual screening workflow works

The featured workflow runs a learned target-aware ranking and structure-based docking as separate evidence branches. Use a method-specific workflow instead when the program lacks a suitable target structure or protein sequence.

  1. Provide target evidence. Submit a compatible protein sequence for SPRINT and a reviewed PDB structure for the structure-based branch.
  2. Run target-aware ranking. Rank the submitted SMILES library with SPRINT and retain its native cosine similarity alongside compound identifiers.
  3. Prepare and dock. Repair the receptor with PDBFixer and dock the library independently with GNINA.
  4. Review independent evidence. Check GNINA geometry with PoseBusters, calculate residue-level interactions with ProLIF, and review descriptors and ADMET predictions.
  5. Compare and shortlist. Compare the branches without treating their raw scores as interchangeable, then export a diverse shortlist for orthogonal review.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Target structure. PDB The main template accepts a PDB receptor. If the source structure is CIF, convert it to PDB before starting the run.
  • Compound library. SMILES Submit a SMILES library directly. Convert existing SDF with SDF to SMILES or MOL2 with MOL2 to SMILES before launching; use the ligand-preparation variant to generate or repair structure files from SMILES.
  • Binding-site context. Known pocket coordinates or a reference ligand can inform manual docking setup and review, but they are not required inputs on this template.

Outputs

  • Prepared target and pocket results. PDB JSON Download the fixed receptor and preparation provenance, plus fpocket pocket files and metrics.
  • Property and rule tables. CSV JSON SMILES Inspect per-compound Lipinski and lead-likeness results with the values and warnings returned by each step.
  • Docking poses and scores. SDF CSV JSON Review ranked docking poses with GNINA CNN score, CNN affinity (pKd), and Vina score (kcal/mol).
  • Pose and ADMET review. CSV JSON Inspect downloadable PoseBusters validation tables and ADMET-AI prediction tables as separate evidence layers.
  • Run record. LOG FILES Retain settings, node-level results, logs, and downloadable files for review or handoff.

Featured virtual screening workflow

SPRINT cosine similarity and GNINA docking outputs remain separate because they have different meanings. PoseBusters and ProLIF review GNINA poses, while molecular descriptors and ADMET-AI provide independent compound-level context.

Hybrid target-aware and structure-based screeningRead-only preview

Inputs

3 required

Methods

7 connected

  1. 01SPRINT Ranking
  2. 02PDBFixer
  3. 03Molecular descriptors
  4. 04GNINA Docking
  5. 05ADMET-AI
  6. 06PoseBusters
  7. 07ProLIF

SPRINT cosine similarity and GNINA docking outputs remain separate because they have different meanings. PoseBusters and ProLIF review GNINA poses, while molecular descriptors and ADMET-AI provide independent compound-level context.

Use this template

Tools for virtual screening

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Start screening