ANARCI icon

ANARCI

(v2026.2.13.2)Code (opens in a new tab)Paper (opens in a new tab)Docs

Number antibody and T cell receptor sequences with multiple numbering schemes

Input

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Run ANARCI”.

What is ANARCI?

ANARCI (Antigen Receptor Numbering And Receptor Classification) is a sequence-analysis program for numbering antibody and T cell receptor (TCR) variable domains. Developed by James Dunbar and Charlotte Deane at the Oxford Protein Informatics Group, it aligns input sequences to Hidden Markov Models built from germline gene databases and maps each residue to a position in the chosen numbering scheme.

Antibody sequences from different organisms and germlines can vary in length, especially around CDR loops. Numbering schemes solve this by defining a universal coordinate system: position 27 in one antibody corresponds to the structurally equivalent position 27 in another. ANARCI automates the assignment across six schemes (IMGT, Chothia, Kabat, Martin, AHo, Wolfguy) while reporting each sequence's chain type, best HMM species match, and optional closest germline gene.

How does ANARCI work?

ANARCI builds one HMM per species and chain type combination using pre-aligned V-gene and J-gene segments from the IMGT/Gene Database. All possible V-J gene combinations form putative germline domain sequences, aligned to MUSCLE with a gap-open penalty of -10. The resulting multiple sequence alignment is converted into a profile HMM using HMMER's hmmbuild with the --hand option to preserve positional structure. The hosted app uses ANARCI's packaged HMM library and supports the source species set: human, mouse, rat, rabbit, rhesus, pig, alpaca, and cow.

When a query sequence arrives, ANARCI runs hmmscan against the full HMM library. The highest-scoring hit determines the chain type (VH, Vκ\kappaκ, Vλ\lambdaλ, Vα\alphaα, Vβ\betaβ, Vγ\gammaγ, Vδ\deltaδ) and species of origin. Alignments scoring below the bit score threshold are rejected, which prevents false recognition of non-immunoglobulin proteins with similar folds. The HMM alignment positions map directly to IMGT numbering; conversion to other schemes applies the insertion and deletion rules defined in each scheme's specification.

In benchmarks on 1.9 million VH sequences from a vaccination study, ANARCI successfully numbered 99.5% of sequences, processing roughly 10,600 sequences per minute on 32 cores.

Numbering schemes

The six supported schemes differ in how they define position equivalence and handle insertions at CDR loops.

SchemeBasisPositionsBest suited for
IMGTGermline gene alignment128 fixed positionsCross-species comparison, standardized reporting
ChothiaStructural alignmentVariableStructure-focused analysis, canonical loop classification
KabatSequence variabilityVariableLegacy datasets, sequence-based CDR definitions
MartinExtended Chothia correctionsVariableStructural engineering with improved indel handling
AHoUnified structural scheme149 fixed positionsBroad structural comparison across domain types
WolfguyAlternative unified schemeVariableSpecialized analyses

Choosing a scheme

IMGT is the most widely adopted for new work. It avoids insertion codes (except in very long CDR3 loops) by assigning each position a single integer from 1 to 128, with unused positions simply skipped. This makes IMGT-numbered sequences straightforward to store in databases and compare computationally.

Chothia and Martin are preferable when structural context matters, since their CDR boundaries align with the physical loop structures observed in crystal structures. Kabat remains important for compatibility with older literature and datasets where CDR definitions are based on sequence variability rather than structure.

AHo uses a fixed 149-position framework that accommodates both antibodies and TCRs under the same numbering, useful for analyses spanning receptor types.

CDR definition differences

The schemes disagree on where CDR loops begin and end. For example, Kabat defines heavy chain CDR1 (HCDR1) starting at position 31, while Chothia starts at position 26 to capture structurally variable residues that Kabat considers framework. IMGT defines all CDRs consistently across chain types: CDR1 at positions 27-38, CDR2 at 56-65, and CDR3 at 105-117. These differences are not cosmetic; the same physical residue can be labeled "CDR" in one scheme and "framework" in another, which affects downstream analyses like humanization scoring or paratope prediction.

How to use ANARCI online

Paste an antibody or TCR protein sequence, or upload FASTA records, then select a numbering scheme and chain types. ProteinIQ runs ANARCI online and returns numbered variable domains, domain boundaries, and HMM match statistics. Optional germline assignment adds the closest V and J genes; native numbering and alignment files are downloadable.

Input

Paste one raw protein sequence, or provide one or more FASTA records. Supported file extensions: .fasta, .fa, .fas, .txt, and gzip-compressed FASTA variants such as .fasta.gz and .fa.gz.

Raw sequences accept the 20 standard amino acid codes and at most 9,999 residues. FASTA files use ANARCI's own parser, which preserves residue case and full headers. The raw-sequence alphabet and length restrictions do not apply to FASTA records. Empty FASTA records follow native ANARCI behavior and can return no numbered domains. File uploads and expanded gzip contents are limited to 50 MB; plan limits still apply.

Leaving the preferred species selection empty searches without a species preference. Combining an empty selection with germline assignment can fail in native ANARCI, as can requesting a chain whose selected species lacks the required germline data.

Settings

Numbering configuration

SettingDescription
Numbering schemeWhich scheme to apply. IMGT (default and recommended), Chothia, Kabat, Martin, AHo, or Wolfguy.
Preferred speciesPrefer species during HMM matching and germline assignment. Default: Human, Mouse. If none of the preferred-species hits clears the bit-score threshold, ANARCI falls back to its full HMM species set. An empty selection applies no species preference; missing germline data can cause a native error when assignment is enabled.
Allowed chain typesRestrict which chain types are matched. Default: all seven (H, K, L, A, B, G, D). Chothia, Kabat, Martin, and Wolfguy can number only immunoglobulin chains; choose IMGT or AHo for TCR-only input.

Advanced options

SettingDescription
Bit score thresholdMinimum HMM alignment score for accepting a hit, including fractional values (ANARCI default 80). Higher values reject more borderline alignments. The original ANARCI paper uses 100; the current ANARCI default accepts slightly more divergent sequences.
Assign germline genesWhen enabled, identifies the closest V and J germline genes for each sequence. Adds processing time but is useful for germline usage analysis and somatic hypermutation studies.

Output columns

ColumnDescription
Query IDSequence identifier from the FASTA header.
DomainOne-based domain number within the query sequence.
Chain TypeIdentified domain: H (heavy), K (kappa), L (lambda), A/B/G/D (TCR alpha/beta/gamma/delta).
SpeciesSpecies of the best-matching HMM. This is a sequence-similarity result, not a definitive biological-origin annotation.
Germline SpeciesSpecies used for the closest germline assignment when germline assignment is enabled.
V GeneClosest V germline gene (when germline assignment is enabled).
V IdentitySequence identity to the closest V germline gene.
J GeneClosest J germline gene (when germline assignment is enabled).
J IdentitySequence identity to the closest J germline gene.
SchemeNumbering scheme applied.
E-valueStatistical significance of the HMM alignment. Lower is better.
Bit ScoreHMM alignment quality score. Higher indicates a stronger match to known immunoglobulin domains.
Numbered SequenceThe variable domain sequence with position assignments in the selected scheme.
Domain Start / Domain EndZero-based, inclusive indices of the first and last numbered residues in the original sequence.

Output files

Each run also returns ANARCI's native files:

FileDescription
anarci_numbering.txtNative vertical residue numbering output.
anarci_*.csvNative per-chain CSV alignments written by ANARCI.
anarci_hits.txtHMMER hit tables for matched domains.
anarci_diagnostics.txtCaptured native ANARCI standard output and error, including species fallback messages.

Interpreting results

A high bit score (typically above 100) with a low E-value indicates confident domain identification and numbering. Sequences scoring near the threshold may represent unusual variants, heavily mutated sequences, or non-immunoglobulin proteins with Ig-like folds.

When a sequence contains multiple domains (e.g., an scFv with both VH and VL), ANARCI reports each domain as a separate row. The Domain Start and Domain End columns indicate where each domain falls in the original sequence.

Input sequence indices and scheme positions are different coordinate systems. For example, Domain Start = 20 and Domain End = 130 select input residues 21 through 131, a span of 111 residues. They do not mean that ANARCI assigned positions 20 through 130. Use the numbered sequence and insertion codes when locating a mutation under IMGT, Kabat, or another scheme.

Species misclassification can occur with highly engineered or chimeric antibodies. ANARCI should not be used as the primary species annotation method. Choosing preferred species guides matching, but ANARCI deliberately falls back to any species when none of those hits clears the bit-score threshold.

Limitations

  • Sequences with unusual insertions or deletions from sequencing errors may fail to number correctly
  • The rigid HMM framework cannot accommodate novel structural features absent from germline databases
  • Species classification reflects germline similarity, not true biological origin, which can mislead for chimeric or heavily engineered sequences
  • PDB coordinate renumbering through ANARCI's separate ImmunoPDB example script is not available in this sequence tool
  • TCR support covers alpha, beta, gamma, and delta chains but with fewer germline references than antibody chains

Table of contents

Related tools

IgBLAST

IgBLAST

Analyze antibody and T cell receptor variable domain sequences

FoldSeek

FoldSeek

Search AlphaFold DB, compare structures, or cluster by 3D similarity

ANARCII

ANARCII

Language-model numbering for antibodies, T cell receptors, and VNAR/VHH domains

BLAST Search

BLAST Search

Find protein sequence matches in Swiss-Prot and PDB.

HMMER

HMMER

Sensitive sequence homology search using profile hidden Markov models

MAFFT

MAFFT

Align protein or nucleotide sequences with selectable accuracy and speed trade-offs.

MMseqs2

MMseqs2

Search and cluster protein or nucleotide sequences for homology discovery at large scale.

MUSCLE5

MUSCLE5

Align multiple protein or nucleotide sequences with high-accuracy PPP refinement.

StringZilla v5

StringZilla v5

Hardware-accelerated edit distances and global or local sequence scores

USAlign

USAlign

Universal structure alignment for proteins, RNA, and DNA molecules