
Structure-based de novo antibody and nanobody design pipeline combining antibody-tuned RFdiffusion, ProteinMPNN sequence design, and antibody-tuned RoseTTAFold2 filtering.

Exploratory antibody CDR co-design for antibody-antigen complexes using Proteo-R1 reasoning and raw diffusion. The standard online workflow does not include the framework structure-inpainting assets required for the published-quality target.

BoltzGen uses generative diffusion models to design protein, peptide, nanobody, and Fab binders against protein and small-molecule targets.

AI-powered antibody CDR design using equivariant diffusion models. Generates complementarity-determining region (CDR) sequences and structures for antibody structures and antibody-antigen complexes. Supports single- and multi-CDR co-design, antibody optimization, fixed-backbone sequence design, and structure prediction.

ProFam-1 is a protein family language model for family-conditioned sequence generation. Provide a protein family in FASTA, A2M, or A3M format and generate new sequences with model likelihood scores for downstream ranking and screening.

EvoDiff is a diffusion-based protein sequence generation framework from Microsoft Research. ProteinIQ currently runs the EvoDiff-Seq OA_DM_38M model for unconditional protein generation, motif scaffolding, and user-sequence inpainting.

Generate protein structures and scaffolds with Genie 3, an all-atom SE(3)-equivariant diffusion model. Genie 3 supports unconditional protein generation, motif scaffolding, and hotspot-targeted binder design.

GenMol is a generative AI model from NVIDIA that creates novel drug-like molecules using masked discrete diffusion. It generates molecules in SAFE representation format and supports de novo generation, linker design, motif extension, and scaffold decoration.

All-atom generative AI for designing protein binders. Specify target binding sites and generate diverse binding proteins with fine-grained control over interaction parameters.

PepMimic designs short peptides that mimic the binding interface of a known protein binder on its target. From a reference protein complex, a latent diffusion model generates peptide candidates constrained to the target interface, and each candidate is scored by interface-mimicry against the reference binder.

Structure-based de novo antibody and nanobody design pipeline combining antibody-tuned RFdiffusion, ProteinMPNN sequence design, and antibody-tuned RoseTTAFold2 filtering.

Exploratory antibody CDR co-design for antibody-antigen complexes using Proteo-R1 reasoning and raw diffusion. The standard online workflow does not include the framework structure-inpainting assets required for the published-quality target.

BoltzGen uses generative diffusion models to design protein, peptide, nanobody, and Fab binders against protein and small-molecule targets.

AI-powered antibody CDR design using equivariant diffusion models. Generates complementarity-determining region (CDR) sequences and structures for antibody structures and antibody-antigen complexes. Supports single- and multi-CDR co-design, antibody optimization, fixed-backbone sequence design, and structure prediction.

ProFam-1 is a protein family language model for family-conditioned sequence generation. Provide a protein family in FASTA, A2M, or A3M format and generate new sequences with model likelihood scores for downstream ranking and screening.

EvoDiff is a diffusion-based protein sequence generation framework from Microsoft Research. ProteinIQ currently runs the EvoDiff-Seq OA_DM_38M model for unconditional protein generation, motif scaffolding, and user-sequence inpainting.

Generate protein structures and scaffolds with Genie 3, an all-atom SE(3)-equivariant diffusion model. Genie 3 supports unconditional protein generation, motif scaffolding, and hotspot-targeted binder design.

GenMol is a generative AI model from NVIDIA that creates novel drug-like molecules using masked discrete diffusion. It generates molecules in SAFE representation format and supports de novo generation, linker design, motif extension, and scaffold decoration.

All-atom generative AI for designing protein binders. Specify target binding sites and generate diverse binding proteins with fine-grained control over interaction parameters.

PepMimic designs short peptides that mimic the binding interface of a known protein binder on its target. From a reference protein complex, a latent diffusion model generates peptide candidates constrained to the target interface, and each candidate is scored by interface-mimicry against the reference binder.
Configure inputs to begin
Set options on the left, then click “Submit job”.
BioPhi is an open-source antibody engineering platform developed by Merck that combines deep learning humanization with repertoire-based humanness evaluation. The platform features two complementary systems trained on the Observed Antibody Space (OAS) database: Sapiens for automated humanization and OASis for humanness scoring.
BioPhi supports sequence-level antibody humanization and comparison against human antibody repertoires. Its outputs are engineering evidence rather than predictions of clinical immunogenicity or retained binding.
BioPhi accepts antibody variable domain sequences in FASTA format. Both heavy (VH) and light (VL) chains can be processed individually or in batches. The platform auto-detects chain types based on sequence characteristics—heavy chains typically start with EVQ/QVQ/DVQ motifs and are longer (~120 residues) than light chains (~110 residues), which often begin with DIQ/EIV/SSE motifs.
>antibody_1
EVQLVESGGGLVQPGGSLRLSCAASGFTFSSYAMSWVRQAPGKGLEWVSAISGSGGSTYYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCAR| Mode | Function |
|---|---|
Sapiens | Generates one final humanized sequence per input chain |
OASis | Evaluates humanness without sequence modification |
Both | Humanizes and evaluates the input and final sequences |
Sapiens positional scores | Returns the source 20-amino-acid score matrix for every position |
Sapiens mean score | Returns one source Sapiens score per input chain |
Sapiens full native report | Returns BioPhi's native alignments, FASTA, and XLSX report |
| Setting | Range | Default | Purpose |
|---|---|---|---|
Humanization iterations | 1–10 | 1 | Number of Sapiens passes applied to each sequence |
Numbering scheme | Kabat/Chothia/IMGT/AHo | Kabat | Numbering scheme passed to BioPhi |
CDR definition | Kabat/Chothia/IMGT/North | Kabat | Defines CDR boundaries |
Humanize CDRs | On/Off | Off | Enables CDR modification |
The default Sapiens run preserves CDRs and applies one humanization iteration. Iterations are successive humanization passes, not a request for multiple alternative designs.
| Setting | Options | Default | Threshold |
|---|---|---|---|
Prevalence threshold | Loose/Relaxed/Medium/Strict | Relaxed | Minimum subject frequency (1%/10%/50%/90%) |
Prevalence threshold controls stringency—stricter thresholds require 9-mer peptides to appear in a higher percentage of human subjects to be considered "human-like."
Results are returned in a spreadsheet with the following columns:
| Column | Description |
|---|---|
Sequence ID | Original input identifier |
Chain type | Heavy or light chain |
Identity % | Sequence identity to original input |
OASis identity (%) | Percentage of evaluated 9-mers classified as human at the selected threshold |
OASis percentile (%) | Percentile relative to BioPhi's reference therapeutic-antibody distribution |
Parental OASis identity (%) | OASis identity for the input sequence |
Parental OASis percentile (%) | OASis percentile for the input sequence |
Mutations | Number of amino acid changes |
Mutation details | Specific substitutions (e.g., "A23G, T45S") |
V germline | Closest V gene match |
J germline | Closest J gene match |
Germline % | Germline content percentage |
Humanized sequence | Output sequence in FASTA format |
Length | Residue count |
The Files tab includes results.csv, submitted sequences, and every native file produced by the selected source mode. These include humanized.fasta for fast Sapiens runs, OASis XLSX workbooks, source score CSV files, or alignments.txt, humanized.fa, and Sapiens.xlsx for a full report.
Sapiens is a BERT-style language model trained on variable domain sequences from 266 human subjects in the OAS database. The model learns probability distributions for each amino acid at each position by training to predict masked or mutated residues in unaligned sequences.
During humanization, Sapiens evaluates the input sequence and computes likelihood scores for all 20 amino acids at every position. Non-human residues—those with low probability in human antibody space—are identified and replaced by sampling from the model's probability distribution. This approach captures complex sequence dependencies that simple germline-matching methods cannot detect.
The model's attention mechanism allows it to recognize context-dependent humanness patterns. A residue considered non-human in one sequence context may be perfectly human in another, depending on surrounding amino acids and structural constraints.
OASis evaluates humanness by extracting all overlapping 9-amino-acid peptides (9-mers) from the input sequence and searching for exact matches in the OAS database. For each 9-mer, the algorithm calculates prevalence—the percentage of human subjects containing that peptide.
OASis identity is the fraction of evaluated peptides classified as human at the selected prevalence threshold. The percentile compares that identity with BioPhi's reference therapeutic-antibody distribution.
The native workbook includes chain-level peptide and germline detail for investigating which sequence regions drive the result.
In a benchmark of 177 therapeutic antibodies, Sapiens generated humanized sequences comparable in quality to those produced by human experts, achieving high humanness scores while maintaining sequence identity. OASis separated human from non-human sequences with high accuracy and showed correlation with clinical immunogenicity data.
OASis percentile is a relative repertoire metric, while OASis identity reports the percentage of evaluated 9-mers that meet the selected prevalence threshold. Compare candidates under the same threshold and retain the native workbook for chain-level context. BioPhi does not define universal clinical pass/fail cutoffs for either metric.
Evaluate sequence identity, mutation locations, OASis metrics, and germline annotations together. No single BioPhi score establishes retained binding or clinical immunogenicity; candidate selection still requires structure, developability, binding, and experimental review.
Complementarity Determining Regions (CDRs) mediate antigen binding and are typically preserved during humanization. The default setting excludes CDRs from modification, changing only framework regions. Enabling CDR humanization may improve humanness scores but risks altering binding specificity or affinity. Such modifications necessitate thorough experimental validation through surface plasmon resonance, ELISA, or functional assays.
BioPhi operates on variable domain sequences only. Full-length antibodies, constant regions, or single-domain antibodies require preprocessing to extract the variable domain.
Sapiens predicts humanness based on sequence patterns in the training data. The model cannot account for factors outside its training scope, such as rare post-translational modifications, unusual structural constraints, or context-specific immunogenicity in particular patient populations.
OASis identifies sequences similar to those in the OAS database but does not predict clinical immunogenicity. Clinical immune responses depend on factors outside this sequence-repertoire comparison, including HLA haplotype, dosing regimen, and epitope presentation.
BioPhi assigns antibody chain context from the input sequence and FASTA records. Engineered variants with atypical sequence characteristics may require preprocessing before submission.