Use case
Ligand-based virtual screening
Compare a compound library with known active ligands using molecular descriptors, fingerprints, pharmacophores, or learned features, then review and export a traceable shortlist.
Inputs
2 required
Methods
5 connected
- 01Molecular descriptors · references
- 02Molecular descriptors · fingerprints
- 03Molecular descriptors · properties
- 04PAINS filter · review
- 05ADMET-AI
ProteinIQ calculates Morgan and MACCS fingerprints but does not yet perform library-versus-reference Tanimoto ranking. This runnable review template accepts the ranked SMILES set from the validated similarity method used by the program and preserves reference and candidate fingerprint evidence for comparison.
Use this templateWhat is ligand-based virtual screening?
Ligand-based virtual screening (LBVS) is a computational method for ranking compounds using evidence from known ligands rather than the three-dimensional structure of a biological target. Common methods compare fingerprints, descriptors, shapes, pharmacophore features, molecular fields, or learned representations. A high rank means similar under the selected representation and metric; it does not prove shared activity or mechanism.
LBVS begins with reference quality. Structures, assay definitions, activity values, stereochemistry, and chemical diversity determine which region of chemical space the search can recognize. Two-dimensional fingerprints are fast and interpretable, while three-dimensional or learned representations may capture different relationships but introduce conformer, training-domain, or model uncertainty.
The method differs from structure-based screening because it does not evaluate complementarity to a receptor pocket or generate a target-specific pose. Similarity values also have no universal activity threshold: a cutoff depends on the fingerprint, metric, references, and library. Activity cliffs make close analogs especially important to inspect rather than automatically promote.
When to use ligand-based virtual screening
- Known active ligands are available. Use one or more experimentally supported reference compounds to search a broader candidate library.
- The target structure is missing or unreliable. Prioritize compounds without relying on a receptor conformation or a defined docking site.
- You need a fast first-pass ranking. Apply inexpensive descriptors or fingerprints before slower structure-based calculations and experiments.
Benefits of ligand-based virtual screening
- Works without a target structure. Known ligands can support screening when a receptor model is unavailable, incomplete, or unsuitable for docking.
- Scales efficiently to large libraries. Two-dimensional fingerprints and compact descriptors are comparatively inexpensive to calculate and compare.
- Supports several kinds of similarity. A program can compare substructures, topological fingerprints, physicochemical properties, pharmacophore features, fields, or shape.
- Makes reference assumptions explicit. Every ranking can retain the reference ligand, representation, metric, and threshold used to produce it.
Primary limitations
- The method inherits bias from known ligands. A narrow or chemically repetitive reference set can steer the search toward familiar scaffolds and miss different active chemotypes.
- Similar molecules can behave differently. Small structural changes can produce activity cliffs, alter selectivity, or change exposure and toxicity.
- Results depend on molecular representation. Different fingerprints, descriptors, conformations, feature definitions, and similarity metrics can rank the same library differently.
- Activity data require careful curation. Mixed assay formats, uncertain measurements, salts, stereochemistry, and inconsistent identifiers can weaken the reference evidence.
- Experimental validation remains necessary. Similarity supports prioritization but does not establish binding, mechanism, potency, selectivity, safety, or efficacy.
Ligand-based virtual screening methods
Each representation encodes a different notion of molecular resemblance. Compare methods on held-out chemistry and retain the representation and parameters behind every score.
- Two-dimensional fingerprints. Compare circular fingerprints such as Morgan or key-based fingerprints such as MACCS with a defined similarity coefficient.
- Descriptor and field methods. Rank candidates by physicochemical profiles or molecular interaction fields when topology alone is not the intended similarity concept.
- Three-dimensional ligand methods. Use shape or pharmacophore feature arrangements when a credible bioactive conformation and adequate conformer coverage are available.
- Learned representations. Apply validated embeddings or predictive models while keeping training provenance, applicability domain, and uncertainty explicit.
Ligand-based virtual screening applications
LBVS can expand around validated hits, retrieve analogs from purchasable libraries, seek scaffold-diverse neighbors with alternative representations, and prefilter a large collection before docking or assays. Multiple references can represent distinct chemotypes or activity profiles.
The method is less suitable when reference activity is uncertain, assays are mixed without harmonization, or the program seeks chemistry outside the domain represented by known ligands. Inactive data and counterexamples are particularly valuable when evaluating whether a ranking enriches the intended phenotype.
How to do ligand-based virtual screening online
ProteinIQ currently provides a review workflow around externally ranked similarity results. It calculates Morgan and MACCS fingerprints for inspection but does not yet perform the library-versus-reference Tanimoto search itself.
- Curate reference evidence. Choose relevant active and inactive ligands, reconcile identifiers and structures, and preserve comparable assay context.
- Standardize the library. Apply one documented policy for salts, charges, tautomers, stereochemistry, duplicates, and invalid inputs.
- Select and validate the representation. Choose fingerprints, descriptors, shape, pharmacophore, or learned features and test them on held-out compounds.
- Rank with the external similarity method. Keep the matched reference, native score, metric, parameters, and aggregation rule for every candidate.
- Review the ranked set in ProteinIQ. Compare fingerprints, molecular properties, alerts, predicted endpoints, chemical neighborhoods, and failures.
- Export a diverse shortlist. Select compounds across credible chemotypes and validate them with orthogonal target-aware evidence and experiments.
How to interpret ligand-similarity rankings
A Tanimoto or other similarity value is meaningful only with its fingerprint and reference set. Do not compare thresholds across representations as though they share one scale, and keep reference-specific scores visible when several actives are used.
Inspect close neighbors for activity cliffs, assay changes, stereochemical differences, reactive groups, and property shifts. The most useful shortlist may include lower-ranked but structurally diverse candidates that test whether the inferred structure–activity relationship generalizes.
How ligand-based virtual screening works
LBVS starts with reference evidence. The quality of the actives, molecular standardization, selected representation, similarity metric, and diversity strategy determine what the ranking means.
- Curate reference ligands. Choose experimentally supported actives, reconcile structures and identifiers, and keep assay context and activity values attached.
- Standardize the library. Normalize structures consistently, preserve compound identifiers, and document how salts, charges, tautomers, and duplicates are handled.
- Calculate molecular representations. Generate the fingerprints, descriptors, pharmacophore features, conformers, or learned embeddings required by the selected method.
- Rank by similarity. Compare candidates with one or more references using a predefined metric and retain reference-specific scores.
- Review and export a diverse shortlist. Inspect neighbors, properties, liabilities, scaffold diversity, and model applicability before selecting compounds for follow-up.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Reference ligands.
SMILESSDFCSVProvide one or more reviewed ligands with stable identifiers and, when available, comparable activity measurements and assay context. - Candidate library.
SMILESSDFUse standardized compounds with stereochemistry and identifiers preserved throughout representation, ranking, and export. - Similarity method. Define the representation, metric, reference aggregation rule, score direction, and shortlist or diversity criteria before screening.
Outputs
- Similarity ranking.
CSVJSONRetain each candidate’s score, matched reference, representation, metric, and rank. - Descriptors or fingerprints.
CSVJSONExport calculated features and failures so the ranking can be reviewed and reproduced. - Diverse shortlist.
SMILESSDFCSVKeep selected structures connected to similarity evidence, property review, alerts, and diversity groups. - Run record.
LOGFILESRetain preparation rules, reference data, method settings, warnings, and exported files.
Tools for ligand-based virtual screening
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

ChEMBL Download
Retrieves ligand structures and activity records for reference-set curation.

PubChem Download
Downloads candidate compounds and reference structures from PubChem.

Molecular descriptors
Calculates physicochemical descriptors, Morgan fingerprints, and MACCS keys.
Open Babel
Converts and standardizes molecular files before representation and ranking.

Ligand fixer
Sanitizes ligand records and prepares consistent three-dimensional structures.

SMILES to SDF
Generates SDF structures when a three-dimensional ligand method is required.

Lipinski's rule of 5
Reports rule-of-five properties for reference and shortlist review.

Lead-likeness filter
Reviews shortlisted compounds against lead-likeness ranges.

PAINS filter
Flags PAINS substructures as a separate liability-review signal.

ADMET-AI
Adds predicted ADMET endpoints after similarity ranking.

Admetica
Provides an independent set of predicted ADMET properties.

eToxPred
Predicts toxicity and synthetic accessibility for shortlisted compounds.
Other small-molecule discovery workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Structure-based virtual screening
Uses a three-dimensional target to generate and score candidate binding poses, often alongside pocket, property, and pose-quality review.
Shape-based virtual screening
Ranks candidate conformers by three-dimensional overlap with one or more reference ligands, optionally including chemical-feature similarity.
Pharmacophore-based virtual screening
Searches for compounds that match a three-dimensional arrangement of interaction features such as donors, acceptors, aromatic regions, hydrophobic regions, and charge centers.
High-throughput virtual screening
Applies staged filtering, batching, and compute planning to screen much larger libraries while keeping throughput and failure handling explicit.
Inverse virtual screening
Evaluates one compound against many potential targets to generate hypotheses about intended targets, off-targets, selectivity, or repurposing opportunities.
Frequently asked questions
Prefer compounds with reliable structures, relevant assay evidence, and comparable activity measurements. Remove duplicates and resolve salts, stereochemistry, and uncertain identifiers before using the set as screening evidence.
The choice depends on the similarity concept and validation set. Morgan fingerprints emphasize local circular environments, while key-based fingerprints encode predefined features. Compare candidate representations on a representative set rather than assuming one is universally best.
There is no universal cutoff. Score distributions depend on the fingerprint, metric, library, and reference set. Choose thresholds using validation compounds and review how enrichment and chemical diversity change across the range.
Keep the score to each reference visible. Maximum similarity, consensus rules, or reference-specific quotas answer different questions, so define the aggregation method before screening and test it against known actives and inactives.
Highly similar molecules can have sharply different activity or selectivity. Inspect close analogs individually and avoid interpreting a high similarity score as a measured biological result.
Public examples include $5,000 for a published 3D shape-similarity campaign across 1.76 million compounds and $20,000 for a listed AI service that explores up to 10 billion compounds for one target.
A basic fingerprint search is far cheaper than a managed 3D shape, pharmacophore, or learned-model project because reference curation, conformer generation, model validation, and expert shortlist review become major parts of the fee.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, and the workflow displays its credit estimate before submission. Reference preparation, execution, and shortlist review can instead be scoped as a done-for-you engagement.
Use held-out actives and appropriate inactives or decoys when available. Review enrichment, scaffold diversity, applicability, and sensitivity to representation and thresholds before selecting compounds for experiments.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.