ESM-C Mutation Scoring icon

ESM-C Mutation Scoring

v3.4.1.post1Code (opens in a new tab)Paper (opens in a new tab)Docs (opens in a new tab)

Score amino acid substitutions with masked protein language models.

Input

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job”.

ESM-C Mutation Scoring webserver overview

ESM-C Mutation Scoring ranks single amino acid substitutions by how strongly a protein language model prefers each replacement to the original token. It masks one position at a time, keeping the rest of the submitted sequence unchanged.

The webserver runs Biohub's MIT-licensed ESMC-300M, ESMC-600M and ESMC-6B checkpoints on ProteinIQ's compute. Every run returns complete mutation scores, residue entropy, ranked substitutions and plots. ESM-C provides embeddings, hidden states and unmasked logits.

Pricing

Runs cost 23 credits per minute with ESMC-300M or ESMC-600M, and 42 credits per minute with ESMC-6B. The minimum reservation equals one minute at the selected rate, not a minimum final charge. Completed runs are charged in proportion to measured runtime, rounded up to a whole credit; unused reserved credits are returned. Set a spending limit before submission.

Inputs

InputAccepted values and limits
Protein sequence(s)Raw single-letter sequences, FASTA text, or sequences fetched from RCSB.
Uploaded file.fasta, .fa, .fas or .txt; up to 10 MiB per file.
Sequence countUp to 50 sequences per job.
Sequence lengthUp to 2,046 sequence tokens per sequence and 20,000 tokens across the job.

Supported tokens are the 20 standard amino acids plus X, B, U, Z, O, . and - gaps, and | chain breaks. Whitespace is removed and letters are converted to uppercase. Positions retain ambiguous amino acids, gaps and chain breaks; unsupported characters are rejected.

Settings

Model settings

ParameterTypeDefaultDescription
Model variantenumESMC-300MESMC-300M, ESMC-600M or ESMC-6B; the 6B model uses a larger GPU.
Batch sizeinteger1Independent single-mask copies processed together; accepts 1 to 4 for 300M and 600M, and only 1 for 6B.

Outputs

Each sequence produces all seven files below. Filenames include the sequence index and label to distinguish repeated labels. A downloadable run.log records the outcome of the whole job. In the arrays, L is the sequence-token count.

FileContents
*_masked_scores.npzFull (L, 64) masked logits and LLRs, entropy, negative fractions, sequence tokens, vocabulary, positions, and model identity.
*_positions.csvEvery position, original token, entropy in bits, and negative fraction.
*_substitutions.csvComplete ranked substitutions with LLRs in natural-log units.
*_mutation_heatmap.svgThe tutorial’s mutation heatmap, including wild-type markers.
*_entropy.svgThe tutorial’s entropy plot.
*_negative_fraction.svgThe tutorial’s negative-fraction plot.
*_mutation_provenance.jsonExact source, model, sequence, runtime, settings, and interpretation.

The Mutation map, Ranked substitutions, Data, Files, and Logs views retain every sequence. Ambiguous amino acids, gaps, and chain-break tokens keep their input positions; their LLR reference is the submitted token.

Understanding results

For position i, ESM-C sees the original sequence with only that position masked. Scores use the model output at token i + 1, accounting for the leading start token. All other positions retain their original sequence.

  • LLR: ln p(substitution | masked sequence) - ln p(wild type | masked sequence). Positive values mean the model prefers the substitution to wild type; negative values mean it prefers wild type. Wild-type scores are zero. These are sequence preferences, not measured changes in activity or stability.
  • Entropy: -sum(p * log2(p + 1e-9)), in bits. Probabilities are normalized across all 64 vocabulary outputs. The 20 standard amino acids are selected only after normalization for the heatmap.
  • Negative fraction: the fraction of the 20 standard amino acids with LLR below zero. The denominator includes the wild type when it is a standard amino acid.
  • Ranked substitutions: all standard amino-acid alternatives except identity substitutions, ordered by decreasing LLR within each sequence. Ties use sequence position, then alphabetical amino-acid order. Positions start at 1.

The heatmap uses the Biohub tutorial’s amino-acid order: descending isoelectric point. Red indicates negative LLR, white zero, and blue positive LLR, on a symmetric scale. Click a cell or use the arrow keys to inspect its score. The entropy chart and exact selected-position values remain available below the map.

These are single substitutions in the original sequence context. Adding their scores does not model epistasis or predict interactions between mutations. No joint multi-mutant prediction is performed. Unmasked logits from ESM-C embeddings are not mutation scores.

Mutation scoring runs on ProteinIQ’s compute using Biohub’s MIT-licensed code and checkpoints. It adapts the tutorial’s API calls to the documented self-hosted model and executes the tutorial’s entropy and LLR functions. The provenance download records the source commit, checkpoint revision, precision, settings, equations, and indexing.

Table of contents

Related tools

ESM-2

ESM-2

Analyze protein sequences and generate embeddings.

sequence-analysisembeddings+3
AbLang-2

AbLang-2

Predict non-germline residues and generate embeddings for paired or unpaired antibody sequences.

sequence-analysisai-powered+5
CANYA

CANYA

Predict protein aggregation nucleation propensity from amino acid sequences.

sequence-analysismachine-learning+5
ESM-C

ESM-C

Generate protein embeddings with ESM-C.

sequence-analysisai-powered+3
Protein-Sol

Protein-Sol

Predict protein solubility from sequence with feature-level and windowed profile outputs.

sequence-analysisempirical+3
pySCA

pySCA

Identify co-evolving residue sectors in protein families using Statistical Coupling Analysis.

sequence-analysiscoevolution-analysis+3
AbLang

AbLang

Restore residues and generate antibody-specific representations or likelihood scores.

sequence-analysisdeep-learning+2
DR-BERT

DR-BERT

Predict intrinsically disordered regions and per-residue disorder probabilities from protein sequences.

sequence-analysisai-powered+3
EvoIF

EvoIF

Score single and multi-site protein mutations with evolutionary and inverse-folding profiles.

protein-analysisproperty-prediction+3
ProstT5

ProstT5

Bidirectional translation between protein sequences and 3Di structural tokens

structure-predictionsequence-analysis+3