
Design sequences for a structure or an aligned ensemble
Input
Caliby webserver overview
Caliby 0.1 designs protein sequences conditioned on one structure or an aligned conformational ensemble. It also scores the sequence in a structure, packs sidechains, generates backbone ensembles with Protpardelle-1c, and offers explicit structure cleaning. Single-structure and ensemble tasks use separate inputs.
This integration targets the Caliby source and its command-line defaults. Native structures, score tables, configuration files, and logs are retained. Caliby and its published model collection are Apache-2.0. Optional AlphaFold2 parameters are provided by DeepMind under CC-BY-4.0; credit the AlphaFold and AlphaFold-Multimer authors when reporting refolding results.
Pricing
Runs cost 27 credits per minute of measured runtime. The minimum reservation is 27 credits, not a minimum final charge. Completed runs are charged in proportion to elapsed time, rounded up to a whole credit; unused reserved credits are returned. Set a spending limit before submission. Each batch job is metered separately.
Inputs
| Input | Formats | Requirement |
|---|---|---|
| Structures | PDB, CIF | Single-structure design, scoring, packing, ensemble generation, or cleaning. Multiple independent structures may be supplied. |
| Structural ensembles | ZIP | Ensemble design or scoring. Each directory contains a named primary structure and its conformers. |
| Residue constraints | CSV | Optional for design tasks. Uses native chain identifiers and label_seq_id positions. |
| Input selection | TXT | Optional native filename list, or directory names for ensembles. |
| Protpardelle sampling configuration | YAML | Optional data-only configuration for ensemble generation with the reviewed cc95 epoch 3490 partial-diffusion model. |
Each uploaded file may be up to 50 MiB, subject to the plan's file limits. Runs have a one-hour execution limit. A ZIP can contain several ensemble directories. For example, protein/protein.cif is the primary structure and protein/decoy_1.cif is another conformation. PDB and CIF conformers are supported. Directory and file names are preserved.
Single-structure design, scoring, packing, and cleaning also accept .pdb.gz and .cif.gz. Compressed files retain their original bytes and filenames. Ensemble generation requires uncompressed PDB or CIF files. Ensemble ZIPs may retain auxiliary metadata such as scaffold_info.csv; Caliby selects the PDB/CIF conformers itself.
Selection TXT files list one exact structure filename or ensemble directory name per line. Uploaded filenames can contain Unicode characters but cannot contain directory paths. The default Protpardelle sampling configuration uses the cc95 epoch 3490 model with 150 rewind steps; an optional YAML file can adjust its supported sampling settings.
Selected conformations must have matching chain identifiers, residue ordering, and residue numbering. Mismatches fail without automatic alignment or renumbering. Ensemble scoring uses the sequence of the first selected conformation: the named primary by default, or the first remaining conformer when primary inclusion is disabled. A named primary file is still required in either case. The native conformer ordering and default maximum of 32 conformers apply.
Residue constraints and symmetry
The CSV requires pdb_key, matching the structure stem or ensemble directory name. For compressed structures, the native stem removes only .gz: protein.pdb.gz uses protein.pdb as its pdb_key. Supported optional columns are:
| Column | Meaning |
|---|---|
fixed_pos_seq | Preserve the sequence at listed positions. |
fixed_pos_scn | Preserve sidechains at listed positions, which must also be included in fixed_pos_seq. |
fixed_pos_override_seq | Set specified residue identities. |
pos_restrict_aatype | Restrict individual positions to specified amino acids. |
symmetry_pos | Tie sequence choices across symmetry-related positions. |
Constraints are passed directly to Caliby. Use its constraint examples for exact CSV syntax. Distilled models do not support fixed sidechain conditioning and reject those requests.
Settings
Unspecified JSON options retain native defaults. false, 0, and null are passed explicitly when supplied. The source validates scientific combinations and supported ranges.
| Parameter | Type | Default | Description |
|---|---|---|---|
| Task | enum | Single-structure design | Selects design, scoring, packing, generation, or cleaning. Ensemble modes have their own inputs. |
| Model | enum | Native default | caliby for design and scoring; caliby_packer_010 for packing. |
| Native run options | object | {} | Mode-specific run options below. |
| Sampling settings | object | {} | Sequence and sidechain sampling options below. |
| AlphaFold2 settings | object | {} | Optional design refolding settings below. |
| Job name | string | optional | Label for identifying the run in job history. |
Design/scoring models are caliby, soluble_caliby, soluble_caliby_v1, caliby_distill, soluble_caliby_distill, and the published caliby_L checkpoint. Packing models are caliby_packer_000, caliby_packer_010, and caliby_packer_030. Custom checkpoint URLs and unreviewed assets are unavailable.
Native run options
| Parameter | Type | Default | Description |
|---|---|---|---|
seed | integer | 0 | Native random seed, except cleaning. |
num_workers | integer | mode-dependent | 2 for design, ensemble scoring, and packing; 4 for single scoring; 8 for generation and cleaning. |
max_num_conformers | integer or null | 32 | Ensemble tasks only. null derives the limit from the first selected ensemble in this version. |
include_primary_conformer | boolean | true | Includes the primary structure in ensemble conditioning. |
save_local_conditionals | boolean | false | Scoring tasks only. Saves native per-position conditional arrays. |
run_self_consistency_eval | boolean | false | Design tasks only. Runs optional AlphaFold2 refolding. |
input_cfg.n_subsample | integer or null | null | Native structure subsampling for single-structure tasks. |
input_cfg.pdb_name_ext | string or null | empty string | Replaces selection-list filename extensions before loading. |
input_cfg.array_id, input_cfg.num_arrays | integer or null | null | Native partitioning of ensemble or generation inputs. |
num_samples_per_pdb | integer | 32 | Ensemble generation only. |
batch_size | integer | 8 | Ensemble generation only. |
Sampling settings
These settings apply to design, scoring, and packing. Nested options are JSON objects, for example {"potts_sampling_cfg":{"potts_temperature":0.01}}.
| Parameter | Type | Default | Description |
|---|---|---|---|
batch_size | integer | mode-dependent | 4 for design and packing; 16 for scoring. |
num_seqs_per_pdb | integer | 16 for design | Number of designed sequences per input. |
num_workers | integer | mode-dependent | Inherits Native run options num_workers unless overridden here. |
verbose | boolean | true | Native sampling diagnostics. |
omit_aas | array or null | null | Amino-acid letters to omit from sequence sampling. |
gaussian_conformers_cfg.n_conformers | integer | 0 | Additional noisy conformers. |
gaussian_conformers_cfg.noise_std | number | 0.0 | Gaussian coordinate noise. |
potts_sampling_cfg.regularization | string or null | LCP | Native Potts regularization. |
potts_sampling_cfg.potts_sweeps | integer | 500 | Number of sampling sweeps. |
potts_sampling_cfg.potts_proposal | enum | dlmc | dlmc or chromatic. |
potts_sampling_cfg.potts_temperature | number | 0.01 | Final sampling temperature. |
potts_sampling_cfg.rejection_step | boolean | false | Native rejection step. |
potts_sampling_cfg.potts_only_cond | boolean | false | Native conditional-only sampling. |
scn_packing_cfg.num_steps | integer | 50 | Sidechain diffusion steps. |
scn_packing_cfg.step_scale | number | 1.5 | Sidechain diffusion step scale. |
ensemble_ignore_res_idx_mismatch may only be false. Alignment checks cannot be disabled.
Optional refolding
AlphaFold2 refolding uses locally installed ColabDesign and model parameters. It does not submit sequences to an external folding or MSA service.
The pinned Caliby refolding path supports single-chain outputs. Multichain structures fail with the native ColabDesign error, including when use_multimer is enabled. That setting selects model parameters; it does not add multichain input support. Multichain sequence design remains available with refolding disabled.
| Parameter | Type | Default | Description |
|---|---|---|---|
num_models | integer | 5 | Models evaluated per sequence. |
sample_models | boolean | true | Native model sampling. |
num_recycles | integer | 3 | Recycling iterations. |
save_best | boolean | true | Retained native configuration field. The pinned Caliby implementation always saves its best-pLDDT prediction. |
use_multimer | boolean | false | Selects AlphaFold-Multimer with the five official multimer-v3 parameter sets. |
Outputs
The Results tab shows the main native table when the task produces one. Refolding contains self-consistency metrics when requested. Files provides downloads of the original structures, tables, arrays, configuration files, and logs.
| Task | Native outputs |
|---|---|
| Design | seq_des_outputs.csv with identifiers, input/designed sequences, U, and structure paths; sampled CIF structures. |
| Scoring | score_outputs.csv with identifiers, sequences, and U; optional local_conditionals/*.npy arrays. |
| Packing | packing_metrics.csv, packed CIF structures, and aligned input structures. |
| Ensemble generation | Protpardelle structures, input metadata, and sampling configuration files. |
| Cleaning | Explicitly cleaned native structures and logs. |
| Refolding | Self-consistency metrics, predictions, and aligned structures. |
| All tasks | Native configuration, standard output/error, and every generated result file. |
U is Caliby's native Potts energy. It is not converted to a probability, confidence, or binding affinity. Local conditional arrays preserve the native token order and shape. Packing scn_rmsd is in angstroms; chi-angle mean absolute errors are in degrees and chi-angle accuracies are fractions. Refolding retains native sc_ca_rmsd, avg_ca_plddt, and tmalign_score values.
The pinned executable source does not implement the README's mentioned score_inputs_csv or save_potts_params options. Those options are unavailable in this version.
Related tools

HyperMPNN
Design protein sequences optimized for thermal stability from backbone structures.

IgDesign
Design antibody CDR loops from antibody-antigen complex structures using inverse folding.

LigandMPNN
Design protein sequences around ligands, metals, and nucleotides for enzyme engineering and binding-site optimization.

ProteinMPNN
Design amino acid sequences for protein backbones with fixed positions, amino acid biases, and sequence diversity controls.

SolubleMPNN
Design sequences with the ProteinMPNN-family model trained on structures from soluble-protein PDB IDs.

AntiFold
Design antibody sequences from structure with AI-powered inverse folding

ESM-IF1
Design protein sequences from 3D backbone structures with controllable sampling diversity.

ProFam
Family-conditioned protein sequence generation from FASTA, A2M, or A3M input

Boltz Sequence Redesign
Redesign chosen residues on a fixed protein structure.

EvoDiff
Generate protein sequences de novo, scaffold fixed motifs, or inpaint selected regions.