CANYA icon

CANYA

0.0.3

Predict protein aggregation nucleation propensity from amino acid sequences. Learn more

Input

FASTA records with protein sequence identifiers.

0/50,000
0 credits

Output

Configure inputs to begin

Set options on the left, then click “Run CANYA”.

What is CANYA?

CANYA (Convolution Attention Network for amYloid Aggregation) predicts amyloid nucleation propensity from protein sequence. The model was trained on aggregation measurements for more than 100,000 random peptides, then tested on independent sequences. It learns short sequence motifs and interactions between them through a compact convolution and attention architecture.

CANYA scores sequence propensity, not the stability of a folded protein or the rate of aggregation under a specific formulation. Because the training assay used surface-accessible 20-residue peptides fused to Sup35N, predictions for regions buried inside a folded protein still need structural and experimental context. Aggrescan3D complements CANYA when a three-dimensional structure is available.

How to use CANYA online

Run CANYA online by submitting one or more protein sequences as FASTA records or a two-column TSV, choosing the published model or 10-model ensemble, and selecting how 20-residue windows should be summarized. ProteinIQ returns a spreadsheet of nucleation scores and the original tab-delimited CANYA result file.

Inputs

Input formatRequirements
FASTAOne or more records with the first > at the very start of the file. Sequences must use uppercase one-letter amino acid codes without internal spaces.
Two-column TSVOne record per line, with sequence ID and amino acid sequence separated by exactly one tab. Do not include a header row.

CANYA predicts the 20 standard amino acids. A stop symbol (*) truncates the sequence at that position. Records containing X or Z are accepted because the CANYA program skips them, but every submission must contain at least one record CANYA can predict. ProteinIQ reports how many records were skipped, and skipped records are absent from the results.

Settings

SettingDescription
ModeDefault model runs the trained instance used for interpretation in the paper. 10-model ensemble averages the ten most interpretable trained instances and also reports their score standard deviation.
SummarizeFor sequences longer than 20 residues, combines overlapping 20-residue window scores using Median (default), Mean, Maximum, or Minimum. Per-window scores keeps every window instead.

The published work recommends the median for longer sequences because it was the most stable summary across the authors' evaluations. Maximum can help locate a strong local hotspot, but it answers a different question from the overall median and is more sensitive to a single high-scoring window.

The 10-model ensemble always returns per-window predictions and standard deviations, regardless of the selected summary.

Results

The columns depend on the chosen mode and summary:

ColumnMeaning
Sequence IDFASTA header or identifier from the first TSV column.
CANYA nucleation scoreSequence-level score after applying the selected summary. Higher values indicate greater predicted amyloid nucleation propensity.
Window sequenceThe 20-residue subsequence scored when per-window output is returned.
PositionLocation of the scored window in the source sequence.
CANYA predictionMean prediction for a window in ensemble output.
Model standard deviationVariation across the ten ensemble models, which reflects model uncertainty rather than experimental error.

The Files tab contains the native _canya.tsv output for downstream analysis. In workflows, score rows and the complete native TSV are available as separate outputs so you can route either representation to the next step.

How CANYA works

CANYA first converts a peptide into a one-hot encoded amino acid matrix. A convolutional layer with 100 filters learns short motifs, then a self-attention layer models positional effects and interactions between those motifs. A 64-unit dense layer feeds a sigmoid output. The complete model has 17,491 parameters.

The original model accepts up to 20 residues at once. ProteinIQ preserves CANYA's sliding-window behavior for longer proteins by scoring every overlapping 20-residue segment and applying the selected summary function.

Interpreting CANYA scores

CANYA scores rank sequences by the pattern learned from its aggregation assay. A higher score supports greater nucleation propensity within that learned sequence context, but the value is not a kinetic rate constant and should not be read as the percentage of protein that will aggregate.

Useful comparisons keep the following factors constant:

  • Summary function: Median, maximum, and minimum scores are not interchangeable.
  • Sequence context: A high-scoring motif may be solvent-accessible in one protein and buried in another.
  • Construct design: Tags, truncations, mutations, and terminal context can change experimental aggregation.
  • Ensemble uncertainty: A large Model standard deviation means the trained models disagree. Such windows deserve less confidence than equally scored windows with low disagreement.

For engineering work, per-window scores are usually the most actionable output because they identify the exact region driving a sequence-level result. Candidate substitutions can then be checked against protein solubility, structural aggregation patches from Aggrescan3D, and experimental expression or aggregation assays.

Table of contents

Related tools

PolyXpert

PolyXpert

Predict low or high antibody polyreactivity from paired VH and VL variable-domain sequences with the source PolyXpert ESM-2 classifier.

antibodytherapeutics+5
Protein-Sol

Protein-Sol

Predict protein solubility from amino acid sequence using the University of Manchester Protein-Sol method.

sequence-analysisempirical+3
Prot2Prop

Prot2Prop

Predict multiple protein developability properties from amino-acid sequences using a multitask ProstT5 adapter.

protein-analysisdeep-learning+5
EvoIF

EvoIF

Score protein mutations with evolutionary profiles from homologous sequences and inverse folding. EvoIF returns a dimensionless log-odds score for each submitted single or multi-site mutation.

protein-analysisproperty-prediction+3
ThermoMPNN

ThermoMPNN

Predict protein thermostability changes (ΔΔG) for point mutations using a graph neural network. Enables computational saturation mutagenesis screening to identify stabilizing mutations.

protein-analysisproperty-prediction+3
Aggrescan3D

Aggrescan3D

Static-mode Aggrescan3D analysis for per-residue aggregation propensity from a single protein structure.

protein-analysisproperty-prediction+3
AllMetal3D

AllMetal3D

Predict metal and water binding sites in protein structures using 3D convolutional neural networks (AllMetal3D + Water3D).

structure-analysisdeep-learning+3
PROPKA 3

PROPKA 3

Predict pKa values of ionizable groups in proteins and protein-ligand complexes from 3D structure. PROPKA calculates environment-driven pKa shifts for standard ionizable residues, terminal groups, and supported ligand atom types.

protein-analysisproperty-prediction+3
SuperWater

SuperWater

Predict protein hydration sites from a structure using a diffusion model with ESM features and a confidence-filtering head.

structure-analysisai-powered+4
TNP

TNP

Profile nanobody developability with the Therapeutic Nanobody Profiler, including CDR geometry, surface hydrophobicity and charge, clinical-reference flags, and predicted structures.

protein-analysisproperty-prediction+4