
ProteusAI
Learn from measured fitness and prioritize protein variants.
Input
ProteusAI webserver overview
ProteusAI 0.1.1 fits regression models to protein sequences and measured experimental fitness. It evaluates the fitted model, then optionally proposes mutations or ranks a supplied candidate library for subsequent experiments.
Every task fits a new model. Inputs and saved results follow workspace access controls. Fitted models and embedding caches remain within the private job and are not available as reusable model downloads.
Pricing
Runs cost 8 credits per minute of measured runtime. The minimum reservation is 8 credits, not a minimum final charge. Completed runs are charged in proportion to elapsed time, rounded up to a whole credit; unused reserved credits are returned. Set a spending limit before submission. Each batch job is metered separately.
Inputs
| Input | Format and requirements |
|---|---|
Measured variants | One .csv file or pasted CSV, required for every task, containing variant names, protein sequences and numerical experimental fitness. |
Candidate variants | One .fasta, .fa or .fas file, or pasted FASTA, required only for Fit model and rank candidates; each record needs a > header. |
| Upload size | Up to 50 MiB per file and 50 MiB combined per submission. |
| Batch input | Up to 10 complete measured datasets; a supplied candidate library is reused for each dataset. |
Column selections must match the CSV headers exactly. Variant names must be text, such as variant_001; columns containing only numbers or booleans are rejected even when values are quoted. ProteusAI performs its native CSV parsing, sequence encoding and padding without automatic sequence repair, deduplication or fitness normalization.
There is no tool-specific fixed row or sequence-length maximum. Runs remain subject to 32 GiB of memory and a 3,500-second computation limit. Dataset size and split choices must leave enough rows for fitting and evaluation; the default split can fail on very small datasets.
Settings
Task and CSV columns
| Parameter | Type | Default | Description |
|---|---|---|---|
Task (task) | enum | search | search: fit and propose mutations; train: fit and evaluate only; predict: fit and rank the candidate FASTA. |
Sequence column (seqs_col) | string | sequence | Exact CSV header containing protein sequences. |
Fitness column (y_col) | string | fitness | Exact CSV header containing numerical experimental measurements. |
Name column (names_col) | string | name | Exact CSV header containing text variant identifiers. |
Model and representation
| Parameter | Type | Default | Description |
|---|---|---|---|
Predictive model (model_type) | enum | rf | Random forest (rf), K-nearest neighbors (knn), support vector regression (svm), ridge regression (ridge) or Gaussian process (gp). |
Sequence representation (representation) | enum | ohe | One-hot encoding (ohe), blosum50, blosum62, vhse, or ESM-2: esm2_8M, esm2_35M, esm2_150M, esm2_650M; esm2 is the 650M alias. |
Representation batch size (batch_size) | integer | 100 | At least 1; controls native representation batches, including mutation search. Smaller batches reduce ESM-2 memory demand. |
The available representations are numerical encodings and the reviewed MIT-licensed ESM-2 checkpoints. Other pretrained models and uploaded fitted estimators are not supported.
Training and evaluation
| Parameter | Type | Default | Description |
|---|---|---|---|
Dataset split (split) | enum | smart | Native stratified 80/10/10 splitting with a random-split fallback (smart), relative weights (ratios), or explicit row membership (custom). |
Training ratio (train_ratio) | number | 0.8 | Nonnegative training weight, used only with ratios. |
Test ratio (test_ratio) | number | 0.1 | Nonnegative test weight, used only with ratios. |
Validation ratio (val_ratio) | number | 0.1 | Nonnegative validation weight, used only with ratios. |
Split row indices (custom_split) | string | conditional | Required with custom: JSON containing nonempty train, test and val arrays of zero-based CSV data-row indices. |
Keep validation separate (grid_search) | boolean | false | Off fits scikit-learn models on training plus validation rows; on keeps validation separate. This setting does not run a hyperparameter search and cannot be enabled for gp. |
K-fold ensemble (ensemble) | boolean | false | Available for rf, knn, svm and ridge; fits the native ensemble and reports variation across its predictions. |
Number of folds (k_folds) | integer | 5 | Used when the ensemble is enabled; at least 2 and no more than the number of fitting rows. |
Native seed (seed) | integer | 42 | Range 0 to 4,294,967,295; controls native splitting and PyTorch operations but does not seed random-forest NumPy randomness. |
Ratio weights must have a positive total. ProteusAI normalizes the weights, rounds down training and test counts, then assigns the remainder to validation. A zero validation weight can therefore still leave validation rows. Empty partitions can fail native training.
For explicit membership, row 0 is the first data row after the header. Each index must exist and may belong to only one partition. Unassigned rows are excluded from fitting and evaluation; repeated indices within the same partition are retained.
Variant search and candidate ranking
| Parameter | Type | Default | Description |
|---|---|---|---|
Proposal objective (optim_problem) | enum | max | Search only: max builds mutation proposals from higher measured fitness; min uses lower measured fitness. Both return descending predicted fitness. |
Mutation attempts (max_eval) | integer | 10000 | Search only, at least 1; returned variants can be fewer because repeated mutation names and unchanged residues are removed. |
Exploration fraction (explore) | number | 0.1 | Search only, from 0 to 1; probability of proposing a random position and amino acid instead of sampling the measured mutation pool. |
Candidate acquisition function (acq_fn) | enum | greedy | Candidate ranking only: greedy prediction (greedy), expected improvement (ei), upper confidence bound (ucb) or random acquisition (random). |
Prediction batch size (prediction_batch_size) | integer | 10000 | Candidate ranking only, at least 1; number of variants in each prediction batch. |
Search uses ProteusAI's native mutation proposals and greedy ranking. Candidate ranking preserves descending acquisition-score order. The built-in 40 synthetic measured variants example selects ridge regression, one-hot encoding and 40 mutation attempts; these differ from the general defaults above.
Outputs
Results shows the selected task's complete result table. Training and test rows shows fitted/evaluation rows, and Evaluation shows the available native metrics and correlation p-values. Files provides downloadable artifacts. Tables preserve ProteusAI's returned row order.
| Download | Contents and availability |
|---|---|
train_data.csv, test_data.csv | Native measured/predicted values and uncertainty for training and test rows. |
val_data.csv | Separate validation rows when produced with Keep validation separate. |
training_results.csv | Returned training/evaluation table with split labels, produced for every task. |
search_results.csv, <model>_<representation>_predictions.csv | Search results and the native prediction file, produced for mutation search. |
candidate_predictions.csv | Scored candidate FASTA records, produced for candidate ranking. |
ranked_variants.fasta | Search or candidate sequences in returned order, available only when FASTA preserves every name and sequence exactly. |
evaluation.json | Available native evaluation metrics. |
provenance.json | Effective settings, input hashes, software/model identities, post-training split membership and result caveats. |
proteusai.log | Private execution log, including diagnostic errors. |
If variant names or sequences cannot be preserved in FASTA, including names with line breaks, the FASTA download is omitted with an explanation. Complete CSV results remain available. Failed runs can retain the log and files from completed stages.
Understanding results
| Field | Meaning |
|---|---|
y_true / y_value | Experimental fitness where available; y_value is the native file header and y_true is used in returned tables. New variants have no measured value. |
y_predicted | Predicted fitness on the scale of the supplied measurements. |
y_sigma | Standard deviation across ensemble predictions or Gaussian-process posterior uncertainty. A single scikit-learn estimator reports zero for new predictions and empty training/test uncertainty values; zero does not establish certainty. |
acq_score | Score used to prioritize candidates; its meaning depends on the selected acquisition function. |
split | Native train, test or val membership label in the returned training table. |
test_r2, val_r2 | Native R² scores; negative values are possible. Ensemble test R² is the mean of individual models' scores. |
test_pearson, val_pearson, test_ken_tau, val_ken_tau | Pearson or Kendall rank correlations, accompanied by their p-values when available. |
calibration, calibration_ratio | Native absolute-error threshold around predictions and fraction within that threshold, where available. Calibration uses test data. |
y_best | Largest observed fitness recorded by the fitted model, used by acquisition functions. |
With separate validation disabled, scikit-learn models fit on training plus validation rows, but their training CSV omits the merged validation rows. Gaussian processes fit training rows only in the supported configuration. Test rows remain held out for fitting, but calibration uses their residuals and is not independent of test evaluation.
ProteusAI 0.1.1 also misaligns names and predictions in the separate validation CSV when a k-fold ensemble is enabled. Those rows must not be used to assess individual validation variants. Result warnings identify this behavior, and the native files are preserved.
Predictions and acquisition scores prioritize experiments; they are not experimental measurements. JSON uses null for nonfinite values, while native CSV values remain unchanged.
Related tools

Genie 3
All-atom SE(3)-equivariant diffusion for unconditional generation, motif scaffolding, and binder design.

Proteo-R1
Design antibody CDRs with Proteo-R1 reasoning, diffusion, and optional framework inpainting.

AntiFold
Design antibody sequences from structure with AI-powered inverse folding

BindCraft
Design de novo protein binders for target surfaces using structure-guided sequence generation.

Boltz Sequence Redesign
Redesign chosen residues on a fixed protein structure.

BoltzProt-1
Protein, peptide, nanobody and antibody binder design.

ESM-IF1
Design protein sequences from 3D backbone structures with controllable sampling diversity.

EvoDiff
Generate protein sequences de novo, scaffold fixed motifs, or inpaint selected regions.

EvoPro
Genetic algorithm-based protein binder optimization using AlphaFold2 and ProteinMPNN

FreeBindCraft
Design de novo protein binders for target surfaces using structure-guided sequence generation.