ProteusAI icon

ProteusAI

v0.1.1Code (opens in a new tab)Docs

Learn from measured fitness and prioritize protein variants.

Input

CSV with sequence, experimental fitness and variant name columns. Select your column names in settings.

Upload file or drag and dropCSV · up to 50 MB

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job”.

ProteusAI webserver overview

ProteusAI 0.1.1 fits regression models to protein sequences and measured experimental fitness. It evaluates the fitted model, then optionally proposes mutations or ranks a supplied candidate library for subsequent experiments.

Every task fits a new model. Inputs and saved results follow workspace access controls. Fitted models and embedding caches remain within the private job and are not available as reusable model downloads.

Pricing

Runs cost 8 credits per minute of measured runtime. The minimum reservation is 8 credits, not a minimum final charge. Completed runs are charged in proportion to elapsed time, rounded up to a whole credit; unused reserved credits are returned. Set a spending limit before submission. Each batch job is metered separately.

Inputs

InputFormat and requirements
Measured variantsOne .csv file or pasted CSV, required for every task, containing variant names, protein sequences and numerical experimental fitness.
Candidate variantsOne .fasta, .fa or .fas file, or pasted FASTA, required only for Fit model and rank candidates; each record needs a > header.
Upload sizeUp to 50 MiB per file and 50 MiB combined per submission.
Batch inputUp to 10 complete measured datasets; a supplied candidate library is reused for each dataset.

Column selections must match the CSV headers exactly. Variant names must be text, such as variant_001; columns containing only numbers or booleans are rejected even when values are quoted. ProteusAI performs its native CSV parsing, sequence encoding and padding without automatic sequence repair, deduplication or fitness normalization.

There is no tool-specific fixed row or sequence-length maximum. Runs remain subject to 32 GiB of memory and a 3,500-second computation limit. Dataset size and split choices must leave enough rows for fitting and evaluation; the default split can fail on very small datasets.

Settings

Task and CSV columns

ParameterTypeDefaultDescription
Task (task)enumsearchsearch: fit and propose mutations; train: fit and evaluate only; predict: fit and rank the candidate FASTA.
Sequence column (seqs_col)stringsequenceExact CSV header containing protein sequences.
Fitness column (y_col)stringfitnessExact CSV header containing numerical experimental measurements.
Name column (names_col)stringnameExact CSV header containing text variant identifiers.

Model and representation

ParameterTypeDefaultDescription
Predictive model (model_type)enumrfRandom forest (rf), K-nearest neighbors (knn), support vector regression (svm), ridge regression (ridge) or Gaussian process (gp).
Sequence representation (representation)enumoheOne-hot encoding (ohe), blosum50, blosum62, vhse, or ESM-2: esm2_8M, esm2_35M, esm2_150M, esm2_650M; esm2 is the 650M alias.
Representation batch size (batch_size)integer100At least 1; controls native representation batches, including mutation search. Smaller batches reduce ESM-2 memory demand.

The available representations are numerical encodings and the reviewed MIT-licensed ESM-2 checkpoints. Other pretrained models and uploaded fitted estimators are not supported.

Training and evaluation

ParameterTypeDefaultDescription
Dataset split (split)enumsmartNative stratified 80/10/10 splitting with a random-split fallback (smart), relative weights (ratios), or explicit row membership (custom).
Training ratio (train_ratio)number0.8Nonnegative training weight, used only with ratios.
Test ratio (test_ratio)number0.1Nonnegative test weight, used only with ratios.
Validation ratio (val_ratio)number0.1Nonnegative validation weight, used only with ratios.
Split row indices (custom_split)stringconditionalRequired with custom: JSON containing nonempty train, test and val arrays of zero-based CSV data-row indices.
Keep validation separate (grid_search)booleanfalseOff fits scikit-learn models on training plus validation rows; on keeps validation separate. This setting does not run a hyperparameter search and cannot be enabled for gp.
K-fold ensemble (ensemble)booleanfalseAvailable for rf, knn, svm and ridge; fits the native ensemble and reports variation across its predictions.
Number of folds (k_folds)integer5Used when the ensemble is enabled; at least 2 and no more than the number of fitting rows.
Native seed (seed)integer42Range 0 to 4,294,967,295; controls native splitting and PyTorch operations but does not seed random-forest NumPy randomness.

Ratio weights must have a positive total. ProteusAI normalizes the weights, rounds down training and test counts, then assigns the remainder to validation. A zero validation weight can therefore still leave validation rows. Empty partitions can fail native training.

For explicit membership, row 0 is the first data row after the header. Each index must exist and may belong to only one partition. Unassigned rows are excluded from fitting and evaluation; repeated indices within the same partition are retained.

Variant search and candidate ranking

ParameterTypeDefaultDescription
Proposal objective (optim_problem)enummaxSearch only: max builds mutation proposals from higher measured fitness; min uses lower measured fitness. Both return descending predicted fitness.
Mutation attempts (max_eval)integer10000Search only, at least 1; returned variants can be fewer because repeated mutation names and unchanged residues are removed.
Exploration fraction (explore)number0.1Search only, from 0 to 1; probability of proposing a random position and amino acid instead of sampling the measured mutation pool.
Candidate acquisition function (acq_fn)enumgreedyCandidate ranking only: greedy prediction (greedy), expected improvement (ei), upper confidence bound (ucb) or random acquisition (random).
Prediction batch size (prediction_batch_size)integer10000Candidate ranking only, at least 1; number of variants in each prediction batch.

Search uses ProteusAI's native mutation proposals and greedy ranking. Candidate ranking preserves descending acquisition-score order. The built-in 40 synthetic measured variants example selects ridge regression, one-hot encoding and 40 mutation attempts; these differ from the general defaults above.

Outputs

Results shows the selected task's complete result table. Training and test rows shows fitted/evaluation rows, and Evaluation shows the available native metrics and correlation p-values. Files provides downloadable artifacts. Tables preserve ProteusAI's returned row order.

DownloadContents and availability
train_data.csv, test_data.csvNative measured/predicted values and uncertainty for training and test rows.
val_data.csvSeparate validation rows when produced with Keep validation separate.
training_results.csvReturned training/evaluation table with split labels, produced for every task.
search_results.csv, <model>_<representation>_predictions.csvSearch results and the native prediction file, produced for mutation search.
candidate_predictions.csvScored candidate FASTA records, produced for candidate ranking.
ranked_variants.fastaSearch or candidate sequences in returned order, available only when FASTA preserves every name and sequence exactly.
evaluation.jsonAvailable native evaluation metrics.
provenance.jsonEffective settings, input hashes, software/model identities, post-training split membership and result caveats.
proteusai.logPrivate execution log, including diagnostic errors.

If variant names or sequences cannot be preserved in FASTA, including names with line breaks, the FASTA download is omitted with an explanation. Complete CSV results remain available. Failed runs can retain the log and files from completed stages.

Understanding results

FieldMeaning
y_true / y_valueExperimental fitness where available; y_value is the native file header and y_true is used in returned tables. New variants have no measured value.
y_predictedPredicted fitness on the scale of the supplied measurements.
y_sigmaStandard deviation across ensemble predictions or Gaussian-process posterior uncertainty. A single scikit-learn estimator reports zero for new predictions and empty training/test uncertainty values; zero does not establish certainty.
acq_scoreScore used to prioritize candidates; its meaning depends on the selected acquisition function.
splitNative train, test or val membership label in the returned training table.
test_r2, val_r2Native R² scores; negative values are possible. Ensemble test R² is the mean of individual models' scores.
test_pearson, val_pearson, test_ken_tau, val_ken_tauPearson or Kendall rank correlations, accompanied by their p-values when available.
calibration, calibration_ratioNative absolute-error threshold around predictions and fraction within that threshold, where available. Calibration uses test data.
y_bestLargest observed fitness recorded by the fitted model, used by acquisition functions.

With separate validation disabled, scikit-learn models fit on training plus validation rows, but their training CSV omits the merged validation rows. Gaussian processes fit training rows only in the supported configuration. Test rows remain held out for fitting, but calibration uses their residuals and is not independent of test evaluation.

ProteusAI 0.1.1 also misaligns names and predictions in the separate validation CSV when a k-fold ensemble is enabled. Those rows must not be used to assess individual validation variants. Result warnings identify this behavior, and the native files are preserved.

Predictions and acquisition scores prioritize experiments; they are not experimental measurements. JSON uses null for nonfinite values, while native CSV values remain unchanged.

Table of contents

Related tools

Genie 3

Genie 3

All-atom SE(3)-equivariant diffusion for unconditional generation, motif scaffolding, and binder design.

protein-designdiffusion-model+5
Proteo-R1

Proteo-R1

Design antibody CDRs with Proteo-R1 reasoning, diffusion, and optional framework inpainting.

protein-designai-powered+5
AntiFold

AntiFold

Design antibody sequences from structure with AI-powered inverse folding

protein-designai-powered+3
BindCraft

BindCraft

Design de novo protein binders for target surfaces using structure-guided sequence generation.

binder-designai-powered+3
Boltz Sequence Redesign

Boltz Sequence Redesign

Redesign chosen residues on a fixed protein structure.

protein-designsequence-design+2
BoltzProt-1

BoltzProt-1

Protein, peptide, nanobody and antibody binder design.

protein-designbinder-design+3
ESM-IF1

ESM-IF1

Design protein sequences from 3D backbone structures with controllable sampling diversity.

sequence-designdeep-learning+2
EvoDiff

EvoDiff

Generate protein sequences de novo, scaffold fixed motifs, or inpaint selected regions.

protein-designai-powered+3
EvoPro

EvoPro

Genetic algorithm-based protein binder optimization using AlphaFold2 and ProteinMPNN

binder-designai-powered+3
FreeBindCraft

FreeBindCraft

Design de novo protein binders for target surfaces using structure-guided sequence generation.

binder-designai-powered+3