
NuCaliby
Joint protein and coding DNA design from protein backbones
Input
NuCaliby webserver overview
NuCaliby 0.1 designs coding DNA and protein sequences from protein backbones using nucleotide-level Potts models. It also supports direct amino-acid design, design conditioned on backbone ensembles, and backbone ensemble generation with Protpardelle-1C. The hosted tool uses the released nucaliby_v1 checkpoint.
The NuCaliby source and bundled NuCaliby checkpoint are MIT licensed. Cite Banaszewski et al. (2026) when reporting results, and the Protpardelle-1C paper for generated ensembles. ViennaRNA-guided results should also credit the ViennaRNA authors and the Institute for Theoretical Chemistry, University of Vienna.
Pricing
NuCaliby uses a fixed upfront price, starting at 10 credits. The quote depends on backbone size, design or conformer count, sampling sweeps, guidance objectives, and ensemble size. The displayed quote is the price for the submitted job; it does not increase with elapsed runtime. Each batch job is priced separately.
ViennaRNA guidance and ensemble generation generally cost more than ordinary sequence design. Custom ensemble sampling configurations can also increase the quote. Runs retain a one-hour execution limit. Older jobs submitted with runtime billing retain their saved billing terms.
Inputs
| Input | Formats | Requirement |
|---|---|---|
| Backbone structures | PDB, ENT, CIF, MMCIF, PDBX, BCIF | Required for both tasks. Add each backbone as a separate input. |
| Backbone ensembles | ZIP | Optional for sequence design. One archive containing target-stem/conformer.pdb or target-stem/conformer.cif. |
| RNA motif files | TXT, FASTA, FA | Optional for sequence design. Reference each uploaded filename in a motif objective. |
| Ensemble sampling configuration | YAML | Optional for ensemble generation. Uses the released cc95 epoch 3490 partial-diffusion model. |
Each uploaded file may be up to 50 MiB, subject to the plan's file limits. Coordinate files use their first model. Sequence design selects author chain IDs; chain A is the default.
For ensemble design, the directory name must match the primary backbone's filename stem. For example, protein.pdb pairs with protein/conformer_1.pdb in the ZIP. Conformers must describe the same protein with compatible residue ordering and length. The primary backbone is included automatically, and ensemble conditioning averages the members' Potts parameters before sampling.
Without a sampling configuration, ensemble generation uses the published 150-step partial-denoising recipe. Custom YAML must retain search_space.models: [[cc95, '3490', sampling_partial_diffusion]], motifs: [null], and ssadj: [null]. Code instantiation, configuration imports, and interpolation are unavailable.
Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
| Task | enum | Design sequences | design or generate_ensemble. |
| Design alphabet | enum | Coding DNA and protein | Design only. nt samples coding DNA and translates it; aa designs protein sequences directly. |
| Guidance objectives | string | [] | Design only. JSON array of objective specification strings, shown when Custom guidance objectives is enabled. |
| Designs per backbone | integer | 1 | Design only. At least one sampled sequence per backbone. |
| Sampling sweeps | integer | 500 | Design only. Number of sampling sweeps. |
| Sampling temperature | number | 0.01 | Design only. Temperature used by sequence sampling. |
| Chains | string | A | Design only. Comma-separated author chain IDs, or all. |
| Generated conformers per backbone | integer | 31 | Ensemble generation only. At least one conformer per input backbone. |
| Random seed | integer | 0 | Seed for the selected task. |
| Job name | string | optional | Label for identifying the run in job history. |
Guidance objectives
The switch is off by default, which uses [] and applies no guidance. Turning it off restores this default while retaining the draft for the next time the switch is enabled. Saved custom values open the editor automatically.
Objectives can be combined in one JSON array, for example:
["stop:weight=100", "tai:weight=50,organism=bl21_de3"]| Objective | Alphabet | Example | Purpose |
|---|---|---|---|
stop | nt | stop:weight=100 | Penalizes stop codons. |
tai | nt | tai:weight=50,organism=bl21_de3 | Guides tRNA adaptation for bl21_de3, e_coli, or s_cerevisiae. |
motif | nt | motif:file=motif.txt,weight=50,beta=2 | Guides an RNA motif using an uploaded filename. Bundled assets/motifs/sira.txt and assets/motifs/hiv1_fse.txt are also available. |
vienna | nt | vienna:weight=10 | Adds an mRNA folding-energy objective using ViennaRNA. |
lcp | aa | lcp | Penalizes low-complexity protein sequences. |
Nucleotide objectives require Coding DNA and protein; lcp requires Amino acids. Guidance is an optimization objective, not a guarantee that every returned sequence satisfies it. Without stop-codon guidance, coding DNA designs may contain stop codons.
Outputs
| Output | Contents |
|---|---|
| Designs table | All 15 native CSV columns in their original order. Available for sequence design. |
designs.csv | Complete design sequences, settings metadata, and scores. |
| Generated conformers | Selectable conformers in the 3D Structure viewer and downloadable PDB files, with their directory names retained. |
stdout.log, stderr.log | Execution messages and diagnostics for both tasks. |
Ensemble generation returns structures rather than designed sequences. Sequence design does not run a separate refolding or experimental validation step. Check the logs if an input backbone has no rows: the native design program can skip an individual failing backbone while continuing with others.
Understanding results
| Field | Meaning |
|---|---|
pdb_id, chain | Backbone identifier and requested chain selection. |
sample, seed | Sample identifier and random seed. |
alphabet, guidance | Design alphabet and applied objective specifications. |
n_members | Number of backbones used for conditioning, including the primary backbone. |
energy | Potts-model energy of the sampled sequence; compare within the same target and configuration. It is not a physical free energy or confidence score. |
recovery | Fraction of compared amino-acid positions matching the primary backbone's native sequence, from 0 to 1. |
gc | Fraction of coding DNA bases that are G or C, from 0 to 1; absent for amino-acid-only designs. |
tai | Coding DNA tRNA adaptation index, evaluated for BL21(DE3), even when guidance uses another organism; absent for amino-acid-only designs. |
has_internal_stop | Native stop-codon indicator; inspect this alongside the translated protein. |
protein, dna, native_aa | Designed protein, coding DNA, and the primary backbone's reference amino-acid sequence. DNA is empty in amino-acid mode. |
The table retains native sequence characters, including * stop characters. These scores describe model outputs and sequence properties; they do not establish folding, expression, or biological activity.
Related tools
Caliby
Design sequences for a structure or an aligned ensemble
HyperMPNN
Design protein sequences optimized for thermal stability from backbone structures.
IgDesign
Design antibody CDR loops from antibody-antigen complex structures using inverse folding.
LigandMPNN
Design protein sequences around ligands, metals, and nucleotides for enzyme engineering and binding-site optimization.
ProteinMPNN
Design amino acid sequences for protein backbones with fixed positions, amino acid biases, and sequence diversity controls.
SolubleMPNN
Design sequences with the ProteinMPNN-family model trained on structures from soluble-protein PDB IDs.
AntiFold
Design antibody sequences from structure with AI-powered inverse folding
ESM-IF1
Design protein sequences from 3D backbone structures with controllable sampling diversity.
ProFam
Family-conditioned protein sequence generation from FASTA, A2M, or A3M input
Boltz Sequence Redesign
Redesign chosen residues on a fixed protein structure.