TL;DR
- Use Both mode for a first humanization run because it returns the Sapiens sequence together with parental and final OASis metrics.
- Enter one complete heavy-chain variable domain in the VH field and one complete light-chain variable domain in the VL field; ProteinIQ pairs them automatically.
- Start with one Sapiens iteration, Kabat numbering and CDR definition, CDR humanization off, and the relaxed 10% OASis threshold.
- Treat OASis identity and percentile as repertoire-based humanness evidence, not as proof of retained binding or low clinical immunogenicity.
BioPhi turns antibody variable-domain sequences into auditable humanization and humanness results. For a first run, open the BioPhi webserver, enter one heavy-chain sequence in the VH field and one light-chain sequence in the VL field, keep Both mode and the default settings, and submit the job. ProteinIQ pairs the two fields automatically. Both mode returns Sapiens-humanized sequences plus OASis metrics for the parental and final chains, making the tradeoff visible in one result.[1]
Start conservatively: one Sapiens iteration, CDR humanization off, and the relaxed OASis threshold. Review every mutation in the numbered alignment, compare parental and final OASis values under the same threshold, then validate binding, structure, developability, and immunogenicity risk with independent methods. A more human-like sequence is not automatically a functional or clinically non-immunogenic antibody.[2][7]
Open BioPhi online
Open the ProteinIQ BioPhi webserver.
Add complete variable domains
Enter complete parental variable domains in the Heavy chain (VH) and
Light chain (VL) fields.
Choose the workflow
Use Both for a first humanization, OASis for evaluation only, or one of the specialized score, complete-report, CDR-grafting, or Designer modes.
Keep a conservative baseline
Start with one iteration, Kabat numbering and CDR definition, CDR humanization off, and the relaxed 10% OASis threshold.
Inspect the evidence
Compare parental and final OASis metrics, identity, numbered mutations, germline annotations, and the native workbook or alignment.
Validate the candidate
Confirm binding and function experimentally, and assess structure, developability, sequence liabilities, and immunogenicity risk separately.
What is BioPhi?
BioPhi is an open-source antibody engineering platform developed at Merck. It combines Sapiens, a deep-learning humanization method, with OASis, an interpretable humanness metric based on exact peptide matches in human antibody repertoires. The original platform also includes CDR grafting and manual sequence design.[2][3]
ProteinIQ runs BioPhi 1.0.11 and Sapiens 1.1.0. The v1.0.11 release moved Sapiens inference to its Hugging Face implementation. The online tool preserves BioPhi's scientific defaults, accepts the source-supported parser inputs in the applicable modes, and returns the native reports alongside a browsable table.[1][4]
BioPhi is most useful when the question is one of these:
- Which framework substitutions does Sapiens propose for a non-human variable domain?
- How human-like are the submitted or designed chains under an OAS-derived peptide metric?
- What happens when parental CDRs are grafted onto selected human V germlines?
- How do explicit, chain-numbered substitutions change the sequence and OASis profile?
- Which amino acids does the Sapiens model support at each position?
It does not predict antigen binding, expression, aggregation, viscosity, specificity, pharmacokinetics, or clinical anti-drug antibody incidence.
How BioPhi works
BioPhi separates sequence humanization, repertoire-based evaluation, and diagnostic score workflows. ProteinIQ returns the native files from each selected path.
Sapiens proposes context-dependent substitutions
Sapiens is an antibody language model trained on human variable-domain sequences from the Observed Antibody Space. During training, residues in unaligned antibody sequences were masked or mutated and the model learned to recover likely human residues from the surrounding sequence context. The published model used repertoires from 266 human subjects.[2][5]
At inference time, Sapiens predicts support for all 20 amino acids at every position. Humanization selects the most probable residue while preserving CDR positions by default. One iteration produces Sapiens*1; each additional iteration applies the process again to the preceding result. Iterations therefore increase humanization depth. They do not generate an independent library of alternative designs.[2]
The model is context-dependent. A residue can be common in one framework context and unusual in another. This is the main distinction from simply replacing every framework residue with the nearest germline residue.
OASis measures exact 9-mer prevalence
OASis splits a variable-domain sequence into every overlapping nine-residue peptide and searches for exact matches in a database built from human antibody repertoires. A 9-mer is considered human when it appears in at least the selected percentage of human subjects. OASis identity is the percentage of evaluated 9-mers that meet that rule.[2]
The four online presets are:
| Threshold | Minimum subject prevalence | Interpretation |
|---|---|---|
| Loose | 1% | Permissive; a peptide can occur in a small fraction of subjects |
| Relaxed | 10% | BioPhi default and a practical baseline for comparisons |
| Medium | 50% | Requires a peptide to occur in at least half of represented subjects |
| Strict | 90% | Highly stringent; only widely observed peptides qualify |
Changing the threshold changes the definition of a qualifying peptide, so OASis identity values from different thresholds should not be compared directly. The percentile supplies reference context for the identity value at the selected threshold, but it is still a humanness ranking rather than a clinical cutoff.
CDR grafting uses a human framework
CDR grafting places the parental CDRs onto BioPhi-selected or user-specified human V germlines. The automatic setting chooses the closest germline. A family such as IGHV3 constrains the search broadly, while a gene such as IGHV3-23 specifies the target more narrowly.[2][3]
The default also restores parental Vernier residues. These framework positions sit beneath or near the antigen-binding loops and can influence loop conformation and affinity even when they do not contact antigen directly. Their importance is why a graft with a more human framework may still require targeted parental backmutations.[6]
Designer applies explicit numbered mutations
Designer takes one antibody input and applies chain-numbered substitutions such as H35:Y, L46:W, or H100A:F. The position uses the selected antibody numbering scheme, and insertion letters are part of the position. Designer is appropriate when mutations have already been chosen from a Sapiens score matrix, a germline comparison, structural review, or experimental evidence.
Choose the right BioPhi mode
| Goal | Recommended mode | Why |
|---|---|---|
| First humanization and before/after comparison | Both | Returns a Sapiens design and parental plus final OASis metrics |
| Humanness analysis without changing sequence | OASis | Preserves the input and produces the native OASis workbook |
| Fast final Sapiens sequence | Sapiens | Returns one final humanized chain per recognized input record |
| Inspect amino-acid support at every position | Sapiens positional scores | Returns the full 20-amino-acid score matrix |
| Rank or summarize model compatibility | Sapiens mean score | Returns one mean Sapiens score per chain |
| Retain BioPhi's detailed native Sapiens report | Sapiens complete report | Returns alignment, FASTA, and XLSX files |
| Use a classical human-germline framework | CDR grafting | Selects or constrains human V germlines and supports Vernier backmutation |
| Apply a predefined mutation plan | Designer mutations | Applies explicit chain-numbered edits in order |
Both mode is the best default because a humanized sequence without the parental reference is harder to evaluate. OASis mode is the safer choice when the goal is characterization rather than design.
Prepare antibody inputs
Use complete variable domains
Submit complete VH and VL variable domains, not full heavy and light chains with signal peptides and constant regions. BioPhi uses antibody numbering and chain recognition, so partial fragments, constant regions, or non-antibody proteins can fail recognition or produce incomplete analyses.
For the normal Sapiens, Both, positional-score, and mean-score workflows, use the two explicit inputs:
Heavy chain (VH)
[complete heavy-chain variable-domain amino-acid sequence]
Light chain (VL)
[complete light-chain variable-domain amino-acid sequence]Each field accepts exactly one sequence. ProteinIQ writes matching _VH and _VL identifiers before running BioPhi, even when the pasted FASTA headers differ. This avoids the source CLI's unpaired-chain warning while preserving separate chain-level result rows.[3]
Match the format to the mode
Sapiens, Both, positional-score, and mean-score modes use the separate protein VH and VL fields. OASis, complete Sapiens report, CDR grafting, and Designer retain one source-native input field and also accept coding-DNA FASTA or PDB input because those workflows use BioPhi's parser for translation, PDB sequence extraction, chain recognition, and pairing. A paired FASTA uploaded in one of those modes still needs matching base identifiers such as candidate_01_VH and candidate_01_VL.[1][3]
Fast protein inputs should contain standard amino-acid letters. X and a terminal stop marker are supported, but alignment gaps are not. Heavy chains can contain at most 144 residues and the light-chain model supports at most 128 residues. Atypical variable domains can still fail BioPhi's antibody numbering step even when their length and alphabet are valid.
Preserve identifiers for source-native files
The separate VH and VL fields do not depend on user-supplied headers. For the single source-native file accepted by parser-backed modes, use a stable clone identifier and do not reuse its base name for unrelated chains. Keep the submitted FASTA with the downloaded result files.
Recommended first-run settings
| Setting | Starting value | Why |
|---|---|---|
| Processing mode | Both | Shows parental and final OASis metrics with the proposed sequence |
| Humanization iterations | 1 | Establishes the smallest Sapiens intervention before deeper passes |
| Numbering scheme | Kabat | BioPhi command-line default |
| CDR definition | Kabat | BioPhi command-line default and the published Sapiens baseline |
| Humanize CDRs | Off | Preserves the antigen-binding loops during the initial design |
| Prevalence threshold | Relaxed, 10% | BioPhi default and a practical baseline for within-project comparison |
Record the numbering scheme and CDR definition with every design. A residue can move between framework and CDR depending on the definition, which changes whether it is protected from Sapiens and how a numbered mutation is interpreted.
Do not increase iterations and enable CDR humanization at the same time in the first comparison. Change one design pressure at a time so the cause of each mutation remains clear.
Interpret the results
Compare parental and final chains first
In Both mode, start with four columns: Identity %, OASis identity (%), Parental OASis identity (%), and the two percentiles. This shows how much sequence changed and whether the OASis profile moved under a fixed threshold.
A useful result is not simply the highest OASis score. Prefer the smallest, structurally plausible mutation set that reaches the project objective while preserving residues supported by binding and functional evidence.
Read OASis identity and percentile separately
OASis identity (%) is an absolute calculation under the chosen threshold: the percentage of evaluated 9-mers classified as human. OASis percentile (%) is relative to BioPhi's therapeutic-antibody reference distribution. Two sequences can have similar identities but different reference context, or a high percentile without any guarantee of low immunogenicity.
Use the native OASis workbook to find the non-human 9-mers and the regions that drive the chain-level value. Overlapping 9-mers mean that one substitution can affect several peptide windows.
Inspect mutation location, not only mutation count
Mutations counts changes between parental and final numbered chains. Mutation details identifies their positions. Review whether each change is in a framework, CDR, Vernier position, VH/VL interface, buried core, or known antigen-contact region. A small number of poorly placed substitutions can be riskier than a larger set of conservative surface substitutions.
CDR grafting and Designer use numbered alignments, so insertions and deletions do not shift every downstream mutation by raw character index. Retain alignments.txt as the audit record.
Use germline fields as context
The V and J germline assignments and germline percentage describe similarity to selected human germline segments. They are useful for understanding a graft or framework choice, but a high germline percentage does not prove folding, affinity, developability, or low immunogenicity.
Treat Sapiens scores as model evidence
The positional-score CSV reports model support across all 20 amino acids at every residue. It is best used to compare alternatives at the same position and in the same sequence context. The mean score compresses that evidence into one chain-level value. Neither output has a universal clinical threshold.
What the BioPhi paper validated
The 2022 study provides three important but bounded validation results:[2]
| Evaluation | Published evidence | Boundary |
|---|---|---|
| Sapiens humanization | Compared with expert humanization across 177 antibodies, including 25 with known parental sequences and 152 with putative parents | Retrospective in silico comparison, not prospective binding or efficacy validation |
| OASis origin classification | The relaxed threshold reported 94.4% accuracy and 0.972 ROC AUC for separating therapeutic antibodies by origin | Classification of curated therapeutic sequences, not a universal project cutoff |
| OASis and reported ADA incidence | Medium OASis identity correlated with ADA incidence across 217 therapeutics at R = -0.53 and R² = 0.28 | Moderate population-level association with substantial unexplained variance |
The study supports BioPhi as a humanization and humanness tool. It does not establish that a particular OASis identity or percentile predicts a patient's immune response. FDA guidance treats immunogenicity assessment as a case-specific, risk-based program influenced by product and patient factors, not as a single sequence score.[7]
Validate a humanized antibody
Use BioPhi as the beginning of an engineering loop:
- Audit the sequence. Confirm chain pairing, numbering, CDR boundaries, mutation positions, motifs, and liabilities.
- Review structural consequences. Predict the paired Fv with ABodyBuilder3, compare parental and final models, and inspect CDRs, Vernier positions, the VH/VL interface, and buried mutations.
- Profile developability. Use TAP2 for structure-informed antibody property context and add project-appropriate aggregation, self-interaction, charge, hydrophobicity, viscosity, and stability assays.
- Confirm binding and function. Measure affinity, kinetics, specificity, competition, and the relevant cellular or biochemical function for parental and designed variants.
- Assess immunogenicity risk. Combine sequence humanness with HLA or T-cell analyses, product attributes, formulation, route, dose, patient population, and an assay strategy appropriate to the development stage.[7]
Maintain more than one candidate when possible. A small panel spanning conservative Sapiens, deeper Sapiens, and carefully reviewed graft variants is usually more informative than committing to the single highest humanness value.
Troubleshooting BioPhi jobs
| Symptom | Likely cause | Action |
|---|---|---|
Variable chain sequence not recognized | DNA was submitted to a protein-only mode, the sequence is incomplete, or antibody numbering cannot classify it | Confirm the molecule type and variable-domain boundaries. Use a parser-backed mode for coding DNA. |
| Pairing warning in a parser-backed mode | VH and VL records inside the single source-native file do not share a recognized base name | Rename records as name_VH and name_VL, or name_HC and name_LC. The separate VH/VL interface pairs chains automatically. |
| Length or model-capacity error | A fast Sapiens input exceeds 144 heavy-chain or 128 light-chain residues | Remove signal peptide or constant-region sequence and submit only the complete variable domain. |
| No recognized chains from PDB | The structure lacks a recognizable antibody variable domain or its coordinates are incomplete | Inspect chain contents and use a variable-domain PDB. Try protein FASTA if the intended sequence is known. |
| Designer rejects the job | More than one antibody input was submitted or the mutation syntax is invalid | Submit one antibody and use entries such as H35:Y or H100A:F. |
| Complete Sapiens report fails while fast Sapiens succeeds | BioPhi 1.0.11 can fail while assembling its native full report when peptide and score tables do not align | Use Sapiens or Both for the sequence, plus OASis and score modes for separate native outputs. |
An older result shows only Warning: | BioPhi wrote the warning header and details to separate output streams, and the older result retained only the header | Re-run the job. Current results reconstruct the complete message and discard empty severity headers. |
An unpaired-chain warning does not mean the individual chain score is empty. It means BioPhi processed one or more records from a source-native file without a recognized partner. The normal two-field interface prevents this by assigning the pair identifiers automatically.
BioPhi compared with related tools
| Tool | Main question | Use it when |
|---|---|---|
| BioPhi | How can a variable domain be humanized, and how human-like is its sequence? | Humanization, OASis evaluation, CDR grafting, or explicit antibody mutations |
| IgBLAST | Which V, D, and J genes and junctions explain an immune-receptor sequence? | Germline annotation and rearrangement analysis rather than design |
| AntiFold | Which antibody sequences are compatible with a fixed backbone? | Structure-conditioned regional sequence design |
| ABodyBuilder3 | What Fv structure is predicted from paired VH and VL sequences? | Reviewing structural consequences after sequence design |
| TAP2 | Are five modeled Fv properties typical relative to clinical-stage therapeutics? | Developability context beyond sequence humanness |
These methods answer different questions. OASis is not a structure predictor, TAP2 is not a humanizer, and a germline assignment is not an immunogenicity result.
Practical checklist
- Complete VH and VL variable domains are present.
- One sequence is entered in each of the separate VH and VL fields.
- The selected mode supports the submitted protein, coding-DNA, or PDB format.
- Numbering scheme and CDR definition are recorded.
- The first Sapiens comparison uses one iteration with CDR humanization off.
- OASis comparisons use the same prevalence threshold.
- Every mutation is reviewed in the numbered alignment.
- Native FASTA, workbook, CSV, and alignment files are retained.
- Parental and designed antibodies are compared structurally and experimentally.
- Humanness is treated as one component of a broader immunogenicity risk assessment.
Frequently asked questions
Which mode should I use first?
Use Both with one protein sequence in each VH and VL field. It returns the proposed Sapiens sequences and shows OASis values before and after humanization.
How many Sapiens iterations should I use?
Start with one. Additional iterations apply Sapiens repeatedly and usually increase the number of changes. Compare successive depths as separate jobs and validate the smallest sufficient mutation set.
Should I enable CDR humanization?
Usually not in the first run. CDR changes can improve a humanness metric but directly affect the antigen-binding loops. Enable them only with a specific design rationale and strong structural and experimental follow-up.
Which OASis threshold is best?
There is no universal best threshold. Relaxed, 10%, is the BioPhi default and a useful baseline. Use stricter or looser settings only for a defined comparison, and never compare raw identity values across different thresholds as though they were on the same scale.
Does a high OASis percentile mean low immunogenicity?
No. It means the sequence ranks highly in BioPhi's therapeutic-antibody reference context at the selected threshold. Clinical immunogenicity depends on patient, product, treatment, and assay factors beyond sequence humanness.[7]
Can BioPhi process coding DNA or PDB files?
Yes, in OASis, complete Sapiens report, CDR grafting, and Designer modes. The fast Sapiens, Both, positional-score, and mean-score modes require protein FASTA.
Why did BioPhi process VH and VL as separate rows?
BioPhi reports chain-level evidence. ProteinIQ associates the VH and VL fields as one antibody before the run, while separate rows preserve the heavy- and light-chain metrics.
What should I download?
Keep the final FASTA, results.csv, the mode-specific XLSX or score CSV, alignments.txt when present, and the normalized input FASTA. Together they preserve the sequence, settings, detailed evidence, and mutation audit trail.
Sources▼
- Use BioPhi Online ProteinIQ · August 25, 2026. https://proteiniq.io/app/biophi
- BioPhi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning mAbs · 2022. https://doi.org/10.1080/19420862.2021.2020203
- BioPhi v1.0.11 source code GitHub (Merck/BioPhi) · August 25, 2026. https://github.com/Merck/BioPhi/tree/v1.0.11
- BioPhi v1.0.11 release GitHub (Merck/BioPhi) · August 25, 2026. https://github.com/Merck/BioPhi/releases/tag/v1.0.11
- Observed Antibody Space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences Protein Science · 2022. https://doi.org/10.1002/pro.4205
- Antibody framework residues affecting the conformation of the hypervariable loops Journal of Molecular Biology · 1992. https://doi.org/10.1016/0022-2836(92)91010-M
- Immunogenicity Assessment for Therapeutic Protein Products U.S. Food and Drug Administration · 2014. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/immunogenicity-assessment-therapeutic-protein-products

Founder and computational chemist, ProteinIQ
Dr. Matic Broz is the founder of ProteinIQ and a computational chemist. He completed a PhD focused on protein structure, molecular dynamics, and neural networks, and writes about structural biology and scientific software.