NuCaliby icon

NuCaliby

(v0.1)Code (opens in a new tab)Paper (opens in a new tab)Docs

Joint protein and coding DNA design from protein backbones

Input

Add each backbone as a separate input. NuCaliby uses the first coordinate model and the selected author chain IDs.

Upload file or drag and dropPDB, ENT, CIF, MMCIF, PDBX, BCIF · up to 50 MB
0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job” — or start from an example.

Crambin backbone: protein sequence design

Ubiquitin: coding DNA with E. coli guidance

Protein G GB1: 31 backbone conformers

NuCaliby webserver overview

NuCaliby 0.1 designs coding DNA and protein sequences from protein backbones using nucleotide-level Potts models. It also supports direct amino-acid design, design conditioned on backbone ensembles, and backbone ensemble generation with Protpardelle-1C. The hosted tool uses the released nucaliby_v1 checkpoint.

The NuCaliby source and bundled NuCaliby checkpoint are MIT licensed. Cite Banaszewski et al. (2026) when reporting results, and the Protpardelle-1C paper for generated ensembles. ViennaRNA-guided results should also credit the ViennaRNA authors and the Institute for Theoretical Chemistry, University of Vienna.

Pricing

NuCaliby uses a fixed upfront price, starting at 10 credits. The quote depends on backbone size, design or conformer count, sampling sweeps, guidance objectives, and ensemble size. The displayed quote is the price for the submitted job; it does not increase with elapsed runtime. Each batch job is priced separately.

ViennaRNA guidance and ensemble generation generally cost more than ordinary sequence design. Custom ensemble sampling configurations can also increase the quote. Runs retain a one-hour execution limit. Older jobs submitted with runtime billing retain their saved billing terms.

Inputs

InputFormatsRequirement
Backbone structuresPDB, ENT, CIF, MMCIF, PDBX, BCIFRequired for both tasks. Add each backbone as a separate input.
Backbone ensemblesZIPOptional for sequence design. One archive containing target-stem/conformer.pdb or target-stem/conformer.cif.
RNA motif filesTXT, FASTA, FAOptional for sequence design. Reference each uploaded filename in a motif objective.
Ensemble sampling configurationYAMLOptional for ensemble generation. Uses the released cc95 epoch 3490 partial-diffusion model.

Each uploaded file may be up to 50 MiB, subject to the plan's file limits. Coordinate files use their first model. Sequence design selects author chain IDs; chain A is the default.

For ensemble design, the directory name must match the primary backbone's filename stem. For example, protein.pdb pairs with protein/conformer_1.pdb in the ZIP. Conformers must describe the same protein with compatible residue ordering and length. The primary backbone is included automatically, and ensemble conditioning averages the members' Potts parameters before sampling.

Without a sampling configuration, ensemble generation uses the published 150-step partial-denoising recipe. Custom YAML must retain search_space.models: [[cc95, '3490', sampling_partial_diffusion]], motifs: [null], and ssadj: [null]. Code instantiation, configuration imports, and interpolation are unavailable.

Settings

ParameterTypeDefaultDescription
TaskenumDesign sequencesdesign or generate_ensemble.
Design alphabetenumCoding DNA and proteinDesign only. nt samples coding DNA and translates it; aa designs protein sequences directly.
Guidance objectivesstring[]Design only. JSON array of objective specification strings, shown when Custom guidance objectives is enabled.
Designs per backboneinteger1Design only. At least one sampled sequence per backbone.
Sampling sweepsinteger500Design only. Number of sampling sweeps.
Sampling temperaturenumber0.01Design only. Temperature used by sequence sampling.
ChainsstringADesign only. Comma-separated author chain IDs, or all.
Generated conformers per backboneinteger31Ensemble generation only. At least one conformer per input backbone.
Random seedinteger0Seed for the selected task.
Job namestringoptionalLabel for identifying the run in job history.

Guidance objectives

The switch is off by default, which uses [] and applies no guidance. Turning it off restores this default while retaining the draft for the next time the switch is enabled. Saved custom values open the editor automatically.

Objectives can be combined in one JSON array, for example:

JSON
["stop:weight=100", "tai:weight=50,organism=bl21_de3"]
ObjectiveAlphabetExamplePurpose
stopntstop:weight=100Penalizes stop codons.
tainttai:weight=50,organism=bl21_de3Guides tRNA adaptation for bl21_de3, e_coli, or s_cerevisiae.
motifntmotif:file=motif.txt,weight=50,beta=2Guides an RNA motif using an uploaded filename. Bundled assets/motifs/sira.txt and assets/motifs/hiv1_fse.txt are also available.
viennantvienna:weight=10Adds an mRNA folding-energy objective using ViennaRNA.
lcpaalcpPenalizes low-complexity protein sequences.

Nucleotide objectives require Coding DNA and protein; lcp requires Amino acids. Guidance is an optimization objective, not a guarantee that every returned sequence satisfies it. Without stop-codon guidance, coding DNA designs may contain stop codons.

Outputs

OutputContents
Designs tableAll 15 native CSV columns in their original order. Available for sequence design.
designs.csvComplete design sequences, settings metadata, and scores.
Generated conformersSelectable conformers in the 3D Structure viewer and downloadable PDB files, with their directory names retained.
stdout.log, stderr.logExecution messages and diagnostics for both tasks.

Ensemble generation returns structures rather than designed sequences. Sequence design does not run a separate refolding or experimental validation step. Check the logs if an input backbone has no rows: the native design program can skip an individual failing backbone while continuing with others.

Understanding results

FieldMeaning
pdb_id, chainBackbone identifier and requested chain selection.
sample, seedSample identifier and random seed.
alphabet, guidanceDesign alphabet and applied objective specifications.
n_membersNumber of backbones used for conditioning, including the primary backbone.
energyPotts-model energy of the sampled sequence; compare within the same target and configuration. It is not a physical free energy or confidence score.
recoveryFraction of compared amino-acid positions matching the primary backbone's native sequence, from 0 to 1.
gcFraction of coding DNA bases that are G or C, from 0 to 1; absent for amino-acid-only designs.
taiCoding DNA tRNA adaptation index, evaluated for BL21(DE3), even when guidance uses another organism; absent for amino-acid-only designs.
has_internal_stopNative stop-codon indicator; inspect this alongside the translated protein.
protein, dna, native_aaDesigned protein, coding DNA, and the primary backbone's reference amino-acid sequence. DNA is empty in amino-acid mode.

The table retains native sequence characters, including * stop characters. These scores describe model outputs and sequence properties; they do not establish folding, expression, or biological activity.

Table of contents

Related tools

Caliby

Caliby

Design sequences for a structure or an aligned ensemble

HyperMPNN

HyperMPNN

Design protein sequences optimized for thermal stability from backbone structures.

IgDesign

IgDesign

Design antibody CDR loops from antibody-antigen complex structures using inverse folding.

LigandMPNN

LigandMPNN

Design protein sequences around ligands, metals, and nucleotides for enzyme engineering and binding-site optimization.

ProteinMPNN

ProteinMPNN

Design amino acid sequences for protein backbones with fixed positions, amino acid biases, and sequence diversity controls.

SolubleMPNN

SolubleMPNN

Design sequences with the ProteinMPNN-family model trained on structures from soluble-protein PDB IDs.

AntiFold

AntiFold

Design antibody sequences from structure with AI-powered inverse folding

ESM-IF1

ESM-IF1

Design protein sequences from 3D backbone structures with controllable sampling diversity.

ProFam

ProFam

Family-conditioned protein sequence generation from FASTA, A2M, or A3M input

Boltz Sequence Redesign

Boltz Sequence Redesign

Redesign chosen residues on a fixed protein structure.