
ESM-C Mutation Scoring
Score amino acid substitutions with masked protein language models.
Input
ESM-C Mutation Scoring webserver overview
ESM-C Mutation Scoring ranks single amino acid substitutions by how strongly a protein language model prefers each replacement to the original token. It masks one position at a time, keeping the rest of the submitted sequence unchanged.
The webserver runs Biohub's MIT-licensed ESMC-300M, ESMC-600M and ESMC-6B checkpoints on ProteinIQ's compute. Every run returns complete mutation scores, residue entropy, ranked substitutions and plots. ESM-C provides embeddings, hidden states and unmasked logits.
Pricing
Runs cost 23 credits per minute with ESMC-300M or ESMC-600M, and 42 credits per minute with ESMC-6B. The minimum reservation equals one minute at the selected rate, not a minimum final charge. Completed runs are charged in proportion to measured runtime, rounded up to a whole credit; unused reserved credits are returned. Set a spending limit before submission.
Inputs
| Input | Accepted values and limits |
|---|---|
Protein sequence(s) | Raw single-letter sequences, FASTA text, or sequences fetched from RCSB. |
| Uploaded file | .fasta, .fa, .fas or .txt; up to 10 MiB per file. |
| Sequence count | Up to 50 sequences per job. |
| Sequence length | Up to 2,046 sequence tokens per sequence and 20,000 tokens across the job. |
Supported tokens are the 20 standard amino acids plus X, B, U, Z, O, . and - gaps, and | chain breaks. Whitespace is removed and letters are converted to uppercase. Positions retain ambiguous amino acids, gaps and chain breaks; unsupported characters are rejected.
Settings
Model settings
| Parameter | Type | Default | Description |
|---|---|---|---|
Model variant | enum | ESMC-300M | ESMC-300M, ESMC-600M or ESMC-6B; the 6B model uses a larger GPU. |
Batch size | integer | 1 | Independent single-mask copies processed together; accepts 1 to 4 for 300M and 600M, and only 1 for 6B. |
Outputs
Each sequence produces all seven files below. Filenames include the sequence index and label to distinguish repeated labels. A downloadable run.log records the outcome of the whole job. In the arrays, L is the sequence-token count.
| File | Contents |
|---|---|
*_masked_scores.npz | Full (L, 64) masked logits and LLRs, entropy, negative fractions, sequence tokens, vocabulary, positions, and model identity. |
*_positions.csv | Every position, original token, entropy in bits, and negative fraction. |
*_substitutions.csv | Complete ranked substitutions with LLRs in natural-log units. |
*_mutation_heatmap.svg | The tutorial’s mutation heatmap, including wild-type markers. |
*_entropy.svg | The tutorial’s entropy plot. |
*_negative_fraction.svg | The tutorial’s negative-fraction plot. |
*_mutation_provenance.json | Exact source, model, sequence, runtime, settings, and interpretation. |
The Mutation map, Ranked substitutions, Data, Files, and Logs views retain every sequence. Ambiguous amino acids, gaps, and chain-break tokens keep their input positions; their LLR reference is the submitted token.
Understanding results
For position i, ESM-C sees the original sequence with only that position masked. Scores use the model output at token i + 1, accounting for the leading start token. All other positions retain their original sequence.
- LLR:
ln p(substitution | masked sequence) - ln p(wild type | masked sequence). Positive values mean the model prefers the substitution to wild type; negative values mean it prefers wild type. Wild-type scores are zero. These are sequence preferences, not measured changes in activity or stability. - Entropy:
-sum(p * log2(p + 1e-9)), in bits. Probabilities are normalized across all 64 vocabulary outputs. The 20 standard amino acids are selected only after normalization for the heatmap. - Negative fraction: the fraction of the 20 standard amino acids with LLR below zero. The denominator includes the wild type when it is a standard amino acid.
- Ranked substitutions: all standard amino-acid alternatives except identity substitutions, ordered by decreasing LLR within each sequence. Ties use sequence position, then alphabetical amino-acid order. Positions start at 1.
The heatmap uses the Biohub tutorial’s amino-acid order: descending isoelectric point. Red indicates negative LLR, white zero, and blue positive LLR, on a symmetric scale. Click a cell or use the arrow keys to inspect its score. The entropy chart and exact selected-position values remain available below the map.
These are single substitutions in the original sequence context. Adding their scores does not model epistasis or predict interactions between mutations. No joint multi-mutant prediction is performed. Unmasked logits from ESM-C embeddings are not mutation scores.
Mutation scoring runs on ProteinIQ’s compute using Biohub’s MIT-licensed code and checkpoints. It adapts the tutorial’s API calls to the documented self-hosted model and executes the tutorial’s entropy and LLR functions. The provenance download records the source commit, checkpoint revision, precision, settings, equations, and indexing.
Related tools

ESM-2
Analyze protein sequences and generate embeddings.

AbLang-2
Predict non-germline residues and generate embeddings for paired or unpaired antibody sequences.

CANYA
Predict protein aggregation nucleation propensity from amino acid sequences.

ESM-C
Generate protein embeddings with ESM-C.

Protein-Sol
Predict protein solubility from sequence with feature-level and windowed profile outputs.

pySCA
Identify co-evolving residue sectors in protein families using Statistical Coupling Analysis.

AbLang
Restore residues and generate antibody-specific representations or likelihood scores.

DR-BERT
Predict intrinsically disordered regions and per-residue disorder probabilities from protein sequences.

EvoIF
Score single and multi-site protein mutations with evolutionary and inverse-folding profiles.

ProstT5
Bidirectional translation between protein sequences and 3Di structural tokens