Use case
High-throughput virtual screening
Prepare and partition a large compound library, apply staged filters, run reproducible batches, and aggregate ranked results without losing failures or method settings.
Inputs
3 required
Methods
6 connected
- 01SPRINT Prefilter
- 02PDBFixer
- 03GNINA Docking
- 04ADMET-AI
- 05PoseBusters
- 06ProLIF
The template uses SPRINT as a target-aware prefilter before GNINA docking. It preserves the native SPRINT and GNINA outputs, checks GNINA poses with PoseBusters and ProLIF, and should be piloted on a representative batch before scaling.
Use this templateWhat is high-throughput virtual screening?
High-throughput virtual screening (HTVS) is a computational execution strategy for evaluating large compound libraries at scale. It combines reproducible preparation, staged filters or rankings, partitioned jobs, failure tracking, aggregation, and focused follow-up. HTVS prioritizes compounds in silico and is distinct from experimental high-throughput screening, which physically measures compounds in assays.
HTVS is not one scoring algorithm. A campaign may use fingerprints, learned target-aware models, shape or pharmacophore filters, fast docking, or a hierarchy of increasingly expensive calculations. The library size, available target or ligand evidence, throughput limits, and desired retained outputs determine the funnel and batch design.
Scale changes the operational risk. Small inconsistencies in identifiers, protonation, software versions, search settings, or score aggregation can affect thousands of records. A defensible run uses deterministic manifests, pilots representative batches, distinguishes invalid compounds from infrastructure failures, and keeps every missing output visible in the final denominator.
When to use high-throughput virtual screening
- The candidate library is too large for one detailed pass. Use staged methods to reserve more expensive calculations for a smaller, better-defined subset.
- Screening must run in reproducible batches. Partition inputs, settings, failures, and outputs so work can be resumed and audited across parallel jobs.
- Throughput and review quality must be balanced. Apply fast early filters without losing the provenance needed for deeper pose and property review.
Benefits of high-throughput virtual screening
- Extends screening to much larger libraries. Batching and parallel execution distribute calculations while preserving a consistent screening protocol.
- Allocates compute by stage. Fast preparation and filtering steps reduce the set that reaches slower docking, rescoring, or predictive methods.
- Makes failures observable. Batch-level records help distinguish invalid compounds, preparation problems, method failures, and infrastructure interruptions.
- Supports reproducible restarts. Stable identifiers, deterministic partitions, and retained settings allow failed or incomplete batches to be rerun without repeating the full screen.
Primary limitations
- Scale multiplies preparation errors. Inconsistent protonation, stereochemistry, identifiers, or malformed structures can propagate across thousands of calculations.
- Fast scoring remains approximate. High-throughput settings trade detail for speed and can increase false positives, false negatives, and unstable rankings.
- Batch effects can distort aggregation. Different settings, software versions, missing outputs, or score normalization choices can make results difficult to compare.
- Compute and storage costs still matter. Large libraries generate substantial intermediate structures, logs, pose files, and result tables even when individual calculations are inexpensive.
- Experimental validation remains necessary. A large computational screen prioritizes compounds for follow-up; it does not establish binding, activity, selectivity, safety, or efficacy.
Large-scale virtual screening
Large-scale virtual screening applies a defined screening protocol to libraries whose size makes serial, high-detail evaluation impractical. The method may use fingerprints, learned models, shape or pharmacophore filters, fast docking, or a funnel that sends progressively smaller subsets to more expensive calculations.
Library size alone does not make a campaign high quality. At large scale, stable identifiers, deterministic batching, comparable settings, explicit failures, retained score provenance, and planned validation become especially important because small preparation or aggregation errors can affect many compounds.
High-throughput screening designs
Choose a design that controls cost without discarding the evidence needed for later review. Thresholds and Top K transitions should be fixed on pilot data before the production run.
- Single-stage parallel screen. Partition one validated method across workers when every compound needs the same calculation and output.
- Hierarchical funnel. Use inexpensive filters or rankings first, then send a documented subset to docking, rescoring, or pose review.
- Parallel evidence branches. Run complementary methods independently when merging too early could hide compounds supported by only one evidence type.
How to do high-throughput virtual screening online
Start with a representative pilot that exercises real library chemistry, expected failures, and the complete output path. Scale only after the manifests and shortlist logic reproduce cleanly.
- Define the campaign budget. Set the library scope, validation controls, batch size, concurrency, retained files, retry policy, and shortlist depth.
- Standardize and inventory the library. Preserve stable identifiers, record normalization rules, and classify duplicates, invalid structures, and excluded chemistry.
- Pilot every stage. Run representative batches through prefiltering, docking or ranking, aggregation, and export to measure failure and output volume.
- Partition deterministic batches. Create manifests that bind compound IDs to method versions, settings, resources, and expected output locations.
- Monitor runs and aggregate explicitly. Separate method failures from infrastructure interruptions, retry only recoverable cases, and never convert missing values into poor scores.
- Rescore and validate the shortlist. Apply slower complementary methods to a focused subset, inspect representative poses, and select diverse compounds for experiments.
How to interpret high-throughput screening results
Compare ranks only across records produced with compatible input preparation, method versions, settings, and score definitions. Batch identity should not predict rank; if it does, investigate configuration drift or aggregation errors before using the shortlist.
Report attrition at every stage, including filters, invalid inputs, timeouts, and missing outputs. Early enrichment, control recovery, scaffold diversity, and representative pose review are more informative than the top score alone.
How high-throughput virtual screening works
HTVS is an execution strategy as much as a scoring method. Library quality, staged filtering, batch design, failure handling, aggregation, and validation determine whether scale produces a useful shortlist.
- Standardize the library. Validate structures, preserve identifiers, define stereochemistry and protonation handling, and remove exact duplicates before partitioning.
- Apply inexpensive early filters. Calculate descriptors and structural alerts, then use predefined inclusion rules to reduce unnecessary downstream calculations.
- Partition and run batches. Create reproducible batches and execute the selected docking or ranking method with consistent versions, settings, and resource limits.
- Aggregate results and failures. Combine batch outputs without silently dropping failed compounds, then keep method-specific scores in their original context.
- Rescore and validate the shortlist. Apply slower or complementary methods to a focused subset and inspect poses, properties, diversity, and known controls.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Prepared compound library.
SMILESSDFUse stable identifiers and a documented standardization policy for salts, charges, stereochemistry, tautomers, and duplicates. - Target or reference evidence.
PDBSDFSMILESProvide the receptor, known ligands, or other query data required by the selected screening method. - Execution plan. Define batch size, concurrency, resource limits, retry behavior, score aggregation, and shortlist criteria before scaling.
Outputs
- Batch manifests and status.
CSVJSONLOGTrack which compounds entered each batch, the settings used, completion state, failures, and retries. - Aggregated ranking.
CSVJSONMerge candidate-level scores while retaining batch, method, version, and failure provenance. - Structures and pose files.
SDFPDBFILESKeep generated structures and selected poses connected to their source compound and ranking record. - Reviewed shortlist.
CSVJSONSDFExport selected candidates with pose checks, properties, alerts, diversity groups, and validation notes.
Tools for high-throughput virtual screening
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

PubChem Download
Downloads compound records and structures for large candidate libraries.

ChEMBL Download
Retrieves known ligands and activity records for controls and reference sets.
Open Babel
Converts and standardizes molecular files before batching.

Molecular descriptors
Calculates descriptors and fingerprints for inexpensive early-stage triage.

Veber's rule
Applies a fast flexibility and polar-surface-area review.

PAINS filter
Flags PAINS substructures before more expensive calculations.

Brenk filter
Reports potentially problematic structural features across the library.

AutoDock-GPU
Runs GPU-accelerated docking for higher-throughput structure-based screens.

SMINA
Provides docking, minimization, and alternative scoring options.

GNINA
Generates ranked poses with CNN and Vina scoring outputs.

PoseBusters
Checks selected docked poses for geometric and chemical plausibility.

ADMET-AI
Adds predicted ADMET endpoints to a focused shortlist.
Other small-molecule discovery workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Structure-based virtual screening
Uses a three-dimensional target to generate and score candidate binding poses, often alongside pocket, property, and pose-quality review.
Shape-based virtual screening
Ranks candidate conformers by three-dimensional overlap with one or more reference ligands, optionally including chemical-feature similarity.
Ligand-based virtual screening
Ranks compounds using similarity, molecular fingerprints, pharmacophores, or learned features derived from known ligands.
Pharmacophore-based virtual screening
Searches for compounds that match a three-dimensional arrangement of interaction features such as donors, acceptors, aromatic regions, hydrophobic regions, and charge centers.
Inverse virtual screening
Evaluates one compound against many potential targets to generate hypotheses about intended targets, off-targets, selectivity, or repurposing opportunities.
Frequently asked questions
HTVS evaluates compounds computationally and produces a prioritized list. Experimental high-throughput screening measures physical samples in assays. A common strategy uses HTVS to reduce the number of compounds selected for experimental testing.
Use deterministic partitions based on stable compound identifiers, keep batch manifests, and apply identical method versions and settings. Batch size should reflect tool limits, expected runtime, failure recovery, and output volume.
They can reduce compute, but aggressive early filters can also remove useful chemistry. Define the scientific rationale and thresholds in advance, retain excluded records, and test the effect on known controls before applying filters to the full library.
Classify failures by cause, such as invalid input, preparation failure, method error, timeout, or infrastructure interruption. Retry only when the cause is recoverable, and keep failures visible in the final denominator and run record.
One current AI screening service publicly lists $20,000 to explore up to 10 billion compounds for one target, including a prioritized hit list, raw data, structures, and a report. At the other end, one university core lists a limited computational screen at $500, illustrating how strongly scope changes the price.
Those services are not directly comparable. Preparation, algorithm, target count, docking or rescoring depth, retained outputs, expert review, compound delivery, and experimental testing can dominate the total beyond raw compute.
ProteinIQ includes self-service HTVS in academic Plus at $29 per month and commercial Pro at $99 per month, and shows the workflow credit estimate before execution. Larger done-for-you campaigns are quoted around batching, compute, failure handling, and deliverables.
Include suitable controls when available, inspect score distributions and failure rates, review representative poses, and test a focused subset with slower orthogonal methods before selecting compounds for experiments.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.