SuperWater icon

SuperWater

1.0.0

Diffusion-based hydration site prediction for protein structures Learn more

Input

Upload files or drag and drop
0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job”.

What is SuperWater?

SuperWater predicts ordered water molecule positions around protein structures. It uses a generative model to sample candidate hydration sites, then filters and clusters those candidates to recover likely crystallographic and interface waters from a single structure.

Published in Communications Chemistry in December 2025, SuperWater was introduced as a score-based diffusion framework with equivariant graph neural networks and ESM-derived features. In the reported benchmarks, it outperformed HydraProt and GalaxyWater-CNN across much of the precision-coverage range for protein surface waters as well as protein-protein and protein-ligand interface waters.

Applications

  • Structure interpretation: Recovering ordered hydration sites that help stabilize local folds and hydrogen-bond networks
  • Binding-site analysis: Highlighting bridging waters near protein-ligand interfaces that may influence affinity and selectivity
  • Protein-protein interfaces: Identifying waters that mediate contacts between chains
  • Model refinement: Adding plausible solvent positions before visualization, inspection, or downstream structural analysis

How to use SuperWater online

ProteinIQ provides browser-based access to SuperWater on hosted compute, so hydration-site prediction can be run from an uploaded structure or an RCSB entry without local setup.

Input

InputDescription
Protein StructureOne protein structure in PDB, ENT, CIF, mmCIF, or PDBx format, or a structure fetched from RCSB. The file must contain protein atoms. Maximum file size is 50 MB.

Settings

SettingDescription
Water ratioApproximate number of sampled water candidates per residue. Range 1-10, default 8. Higher values increase coverage but also runtime and memory use. Larger structures may automatically route to higher GPU tiers, and combinations beyond the supported GPU ceiling are blocked before launch.
Confidence cutoffMinimum confidence score required to keep a predicted hydration site. Range 0.02-0.5, default 0.1. Higher values improve precision but return fewer waters.
Inference stepsReverse-diffusion steps used during sampling. 10 is faster, 20 is the default, and 30 may improve quality at the cost of runtime.

Output

ProteinIQ returns a 3D viewer, a tabular summary, and downloadable result files.

OutputDescription
ViewerInteractive structure view with predicted water coordinates included in the returned files.
Data tableSummary table for the processed structure and prediction settings.
FilesDownloadable structure and result files, including the combined structure with predicted waters and auxiliary confidence outputs.

Output columns

ColumnDescription
StructureInput structure name or identifier
Predicted watersNumber of hydration sites retained after confidence filtering and clustering
Sampled candidatesNumber of candidate positions generated before final filtering
Max candidate confidenceHighest confidence among the sampled candidate positions before clustering
Mean candidate confidenceAverage confidence across all sampled candidate positions before clustering
Water ratioWater sampling multiplier used for the run
Confidence cutoffAcceptance threshold used to keep predicted waters
Input sourceWhether the structure came from file upload or RCSB fetch

How does SuperWater work?

SuperWater does not score a fixed 3D voxel grid the way earlier hydration predictors do. Instead, it learns the gradient of the water-position distribution around a protein and uses that learned score field to refine randomly initialized water coordinates through reverse diffusion.

The architecture reported in the paper combines a score-based diffusion model with equivariant graph neural networks, which preserves geometric consistency under rotation. The published method also incorporates ESM features to encode sequence-derived context around residues. After sampling, a separate confidence model removes low-probability positions, and a clustering step consolidates nearby candidates into final hydration sites.

The benchmark protocol described in the paper sampled an initial number of candidates proportional to protein length, then traced different precision-coverage tradeoffs by varying the internal confidence threshold (cap). On an independent test set of 1,709 crystal structures, SuperWater defined the best overall precision-coverage frontier among the compared methods across many operating points.

Interpreting results

SuperWater predictions represent likely ordered water sites, not every transient solvent molecule around the protein. High-confidence predictions are more likely to correspond to persistent hydration sites that recur in experimental structures or stabilize interfaces.

Confidence patternInterpretation
High Max candidate confidence with moderate-to-high Predicted watersStrong evidence for a structured hydration pattern around the input fold or interface
Low Confidence cutoff and many returned watersBroader coverage, useful for exploratory analysis but more likely to include false positives
Higher Confidence cutoff and fewer returned watersMore conservative set of hydration sites, typically better for prioritizing inspection
Large gap between Sampled candidates and Predicted watersMany sampled positions were rejected during filtering, indicating a stricter or more selective final result

The paper reports mean absolute deviation of approximately 0.3 ± 0.06 Å at cap = 0.5 for matched predictions, which indicates that true-positive sites can be placed with sub-angstrom accuracy. That figure should be interpreted as benchmarked localization accuracy against crystallographic waters, not as a guarantee for every structure.

Limitations

SuperWater predicts positions from a static protein structure. It does not explicitly model long-timescale solvent dynamics, alternate conformations, protonation-state uncertainty, or experimental conditions such as crystal packing, buffer composition, or ligand occupancy.

Like other hydration-site predictors, it is biased toward ordered waters that are recoverable from structural datasets. Disordered, low-occupancy, or rapidly exchanging solvent molecules may be absent from the output even when they are biologically relevant.

Prediction counts also depend directly on Water ratio and Confidence cutoff. A denser sampling strategy can improve recall, but it does not by itself validate the biological importance of every returned site.

ProteinIQ accepts large structures by routing them across multiple GPU tiers, but there is still a hard ceiling for very large residue count × water ratio combinations. If a submission exceeds that supported range, lower Water ratio or use a smaller structure.

Table of contents

Related tools

AllMetal3D

AllMetal3D

Predict metal and water binding sites in protein structures using 3D convolutional neural networks (AllMetal3D + Water3D).

structure-analysisdeep-learning+3
Aggrescan3D

Aggrescan3D

Static-mode Aggrescan3D analysis for per-residue aggregation propensity from a single protein structure.

protein-analysisproperty-prediction+3
PROPKA 3

PROPKA 3

Predict pKa values of ionizable groups in proteins and protein-ligand complexes from 3D structure. PROPKA calculates environment-driven pKa shifts for standard ionizable residues, terminal groups, and supported ligand atom types.

protein-analysisproperty-prediction+3
CANYA

CANYA

Predict protein aggregation nucleation propensity from amino acid sequences using the Lehner Lab CANYA neural network.

sequence-analysismachine-learning+5
EvoIF

EvoIF

Score protein mutations with evolutionary profiles from homologous sequences and inverse folding. EvoIF returns a dimensionless log-odds score for each submitted single or multi-site mutation.

protein-analysisproperty-prediction+3
LocScale

LocScale

LocScale performs physics-informed local sharpening of cryo-EM density maps using half-maps or full MRC/MAP volumes, with optional mask and reference-map inputs.

structure-analysisphysics-based+2
MolProbity

MolProbity

Validate protein structure quality with all-atom contact analysis, Ramachandran plots, rotamer assessment, and geometry checks.

structure-analysisquality-validation+4
PDBsum

PDBsum

Generate a downloadable PDBsum structural summary report archive for a single protein structure.

structure-analysisquality-validation+3
PolyXpert

PolyXpert

Predict low or high antibody polyreactivity from paired VH and VL variable-domain sequences with the source PolyXpert ESM-2 classifier.

antibodytherapeutics+5
Prot2Prop

Prot2Prop

Predict multiple protein developability properties from amino-acid sequences using a multitask ProstT5 adapter.

protein-analysisdeep-learning+5