Use case

High-throughput virtual screening

Prepare and partition a large compound library, apply staged filters, run reproducible batches, and aggregate ranked results without losing failures or method settings.

Tiered high-throughput virtual screeningRead-only preview

Inputs

3 required

Methods

6 connected

  1. 01SPRINT Prefilter
  2. 02PDBFixer
  3. 03GNINA Docking
  4. 04ADMET-AI
  5. 05PoseBusters
  6. 06ProLIF

The template uses SPRINT as a target-aware prefilter before GNINA docking. It preserves the native SPRINT and GNINA outputs, checks GNINA poses with PoseBusters and ProLIF, and should be piloted on a representative batch before scaling.

Use this template

What is high-throughput virtual screening?

High-throughput virtual screening (HTVS) is a computational execution strategy for evaluating large compound libraries at scale. It combines reproducible preparation, staged filters or rankings, partitioned jobs, failure tracking, aggregation, and focused follow-up. HTVS prioritizes compounds in silico and is distinct from experimental high-throughput screening, which physically measures compounds in assays.

HTVS is not one scoring algorithm. A campaign may use fingerprints, learned target-aware models, shape or pharmacophore filters, fast docking, or a hierarchy of increasingly expensive calculations. The library size, available target or ligand evidence, throughput limits, and desired retained outputs determine the funnel and batch design.

Scale changes the operational risk. Small inconsistencies in identifiers, protonation, software versions, search settings, or score aggregation can affect thousands of records. A defensible run uses deterministic manifests, pilots representative batches, distinguishes invalid compounds from infrastructure failures, and keeps every missing output visible in the final denominator.

When to use high-throughput virtual screening

  • The candidate library is too large for one detailed pass. Use staged methods to reserve more expensive calculations for a smaller, better-defined subset.
  • Screening must run in reproducible batches. Partition inputs, settings, failures, and outputs so work can be resumed and audited across parallel jobs.
  • Throughput and review quality must be balanced. Apply fast early filters without losing the provenance needed for deeper pose and property review.

Benefits of high-throughput virtual screening

  • Extends screening to much larger libraries. Batching and parallel execution distribute calculations while preserving a consistent screening protocol.
  • Allocates compute by stage. Fast preparation and filtering steps reduce the set that reaches slower docking, rescoring, or predictive methods.
  • Makes failures observable. Batch-level records help distinguish invalid compounds, preparation problems, method failures, and infrastructure interruptions.
  • Supports reproducible restarts. Stable identifiers, deterministic partitions, and retained settings allow failed or incomplete batches to be rerun without repeating the full screen.

Primary limitations

  • Scale multiplies preparation errors. Inconsistent protonation, stereochemistry, identifiers, or malformed structures can propagate across thousands of calculations.
  • Fast scoring remains approximate. High-throughput settings trade detail for speed and can increase false positives, false negatives, and unstable rankings.
  • Batch effects can distort aggregation. Different settings, software versions, missing outputs, or score normalization choices can make results difficult to compare.
  • Compute and storage costs still matter. Large libraries generate substantial intermediate structures, logs, pose files, and result tables even when individual calculations are inexpensive.
  • Experimental validation remains necessary. A large computational screen prioritizes compounds for follow-up; it does not establish binding, activity, selectivity, safety, or efficacy.

Large-scale virtual screening

Large-scale virtual screening applies a defined screening protocol to libraries whose size makes serial, high-detail evaluation impractical. The method may use fingerprints, learned models, shape or pharmacophore filters, fast docking, or a funnel that sends progressively smaller subsets to more expensive calculations.

Library size alone does not make a campaign high quality. At large scale, stable identifiers, deterministic batching, comparable settings, explicit failures, retained score provenance, and planned validation become especially important because small preparation or aggregation errors can affect many compounds.

High-throughput screening designs

Choose a design that controls cost without discarding the evidence needed for later review. Thresholds and Top K transitions should be fixed on pilot data before the production run.

  • Single-stage parallel screen. Partition one validated method across workers when every compound needs the same calculation and output.
  • Hierarchical funnel. Use inexpensive filters or rankings first, then send a documented subset to docking, rescoring, or pose review.
  • Parallel evidence branches. Run complementary methods independently when merging too early could hide compounds supported by only one evidence type.

How to do high-throughput virtual screening online

Start with a representative pilot that exercises real library chemistry, expected failures, and the complete output path. Scale only after the manifests and shortlist logic reproduce cleanly.

  1. Define the campaign budget. Set the library scope, validation controls, batch size, concurrency, retained files, retry policy, and shortlist depth.
  2. Standardize and inventory the library. Preserve stable identifiers, record normalization rules, and classify duplicates, invalid structures, and excluded chemistry.
  3. Pilot every stage. Run representative batches through prefiltering, docking or ranking, aggregation, and export to measure failure and output volume.
  4. Partition deterministic batches. Create manifests that bind compound IDs to method versions, settings, resources, and expected output locations.
  5. Monitor runs and aggregate explicitly. Separate method failures from infrastructure interruptions, retry only recoverable cases, and never convert missing values into poor scores.
  6. Rescore and validate the shortlist. Apply slower complementary methods to a focused subset, inspect representative poses, and select diverse compounds for experiments.

How to interpret high-throughput screening results

Compare ranks only across records produced with compatible input preparation, method versions, settings, and score definitions. Batch identity should not predict rank; if it does, investigate configuration drift or aggregation errors before using the shortlist.

Report attrition at every stage, including filters, invalid inputs, timeouts, and missing outputs. Early enrichment, control recovery, scaffold diversity, and representative pose review are more informative than the top score alone.

How high-throughput virtual screening works

HTVS is an execution strategy as much as a scoring method. Library quality, staged filtering, batch design, failure handling, aggregation, and validation determine whether scale produces a useful shortlist.

  1. Standardize the library. Validate structures, preserve identifiers, define stereochemistry and protonation handling, and remove exact duplicates before partitioning.
  2. Apply inexpensive early filters. Calculate descriptors and structural alerts, then use predefined inclusion rules to reduce unnecessary downstream calculations.
  3. Partition and run batches. Create reproducible batches and execute the selected docking or ranking method with consistent versions, settings, and resource limits.
  4. Aggregate results and failures. Combine batch outputs without silently dropping failed compounds, then keep method-specific scores in their original context.
  5. Rescore and validate the shortlist. Apply slower or complementary methods to a focused subset and inspect poses, properties, diversity, and known controls.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Prepared compound library. SMILES SDF Use stable identifiers and a documented standardization policy for salts, charges, stereochemistry, tautomers, and duplicates.
  • Target or reference evidence. PDB SDF SMILES Provide the receptor, known ligands, or other query data required by the selected screening method.
  • Execution plan. Define batch size, concurrency, resource limits, retry behavior, score aggregation, and shortlist criteria before scaling.

Outputs

  • Batch manifests and status. CSV JSON LOG Track which compounds entered each batch, the settings used, completion state, failures, and retries.
  • Aggregated ranking. CSV JSON Merge candidate-level scores while retaining batch, method, version, and failure provenance.
  • Structures and pose files. SDF PDB FILES Keep generated structures and selected poses connected to their source compound and ranking record.
  • Reviewed shortlist. CSV JSON SDF Export selected candidates with pose checks, properties, alerts, diversity groups, and validation notes.

Tools for high-throughput virtual screening

Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open screening workflow