Protein-Sol icon

Protein-Sol

(2017-10)

Predict protein solubility from sequence with feature-level and windowed profile outputs. Learn more

Input

Inputs

0/25,000
0 credits

Output

Configure inputs to begin

Set options on the left, then click “Predict solubility”.

What is Protein-Sol?

Protein-Sol is an empirical sequence-based method for predicting protein solubility. It reports the core quantities described by the University of Manchester method: percent-sol, scaled-sol, population-sol, and predicted pI.

The method calculates sequence features such as charge balance, amino acid composition, Kyte-Doolittle hydropathy, fold propensity, disorder propensity, entropy, and beta propensity. These features are compared with the Niwa cell-free expression solubility dataset to estimate solubility.

How to use Protein-Sol online

Paste or upload protein FASTA records to run Protein-Sol online. ProteinIQ checks the identifiers and sequence lengths, runs every record with at least 21 standard amino acids, and reports shorter records as skipped. Results include a prediction table plus downloadable feature, profile, composition, and native run files.

Inputs

InputRequirement
Protein sequencesFASTA text or a .fasta, .fa, .fas, or .txt file.
Sequence lengthAt least one record must contain 21 standard amino acids. Shorter records are skipped and listed in the result warnings.
Sequence identifierThe first whitespace-delimited token after > must be unique among records that are processed.
Job sizeUp to 25 KB. Record limits range from 5 for guests to 100 on the Pro plan.

Use the 20 standard one-letter amino acid codes when possible:

Text
>my_protein
MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQLR

Protein-Sol ignores stop symbols and characters outside the standard amino acid alphabet during its sequence preparation step. For clean, reproducible predictions, submit canonical protein FASTA records.

Results

The Predictions view contains one row per accepted sequence. The same rows are available in the predictions CSV.

ColumnMeaning
IDEffective FASTA identifier reported by Protein-Sol.
percent-solPredicted solubility percentage.
scaled-solPrediction scaled over the source reference range.
population-solReference population solubility on the same scale.
pIPredicted isoelectric point used by the Protein-Sol feature model.

The CSV files view includes downloads for feature weights, feature deviations, sliding-window profiles, and whole-sequence or windowed composition data.

Downloadable Files

Protein-Sol returns CSV tables for the parsed results:

  • predictions.csv: solubility predictions and predicted pI.
  • feature_weights.csv: feature contribution weights used by the model.
  • feature_deviations.csv: sequence feature deviations.
  • profiles.csv: 21-residue window profile values.
  • composition.csv: whole-sequence and windowed composition features.
  • composition_summaries.csv: composition summary lines.

Generated run files are retained with the result download bundle:

  • seq_prediction.txt: prediction rows, feature weights, deviations, and 21-residue profiles.
  • seq_composition.txt: whole-sequence, 21-residue, and 51-residue composition features.
  • run.log: execution log.
  • messages.txt: parser messages when Protein-Sol reports skipped or adjusted input records.

Citation

Hebditch M, Carballo-Amador MA, Charonis S, Curtis R, Warwicker J. Protein-Sol: a web tool for predicting protein solubility from sequence. Bioinformatics. 2017;33(19):3098-3100. doi:10.1093/bioinformatics/btx345

Table of contents

Related tools

CANYA

CANYA

Predict protein aggregation nucleation propensity from amino acid sequences using the Lehner Lab CANYA neural network.

sequence-analysismachine-learning+5
EvoIF

EvoIF

Score protein mutations with evolutionary profiles from homologous sequences and inverse folding. EvoIF returns a dimensionless log-odds score for each submitted single or multi-site mutation.

protein-analysisproperty-prediction+3
PolyXpert

PolyXpert

Predict low or high antibody polyreactivity from paired VH and VL variable-domain sequences with the source PolyXpert ESM-2 classifier.

antibodytherapeutics+5
Prot2Prop

Prot2Prop

Predict multiple protein developability properties from amino-acid sequences using a multitask ProstT5 adapter.

protein-analysisdeep-learning+5
ThermoMPNN

ThermoMPNN

Predict protein thermostability changes (ΔΔG) for point mutations using a graph neural network. Enables computational saturation mutagenesis screening to identify stabilizing mutations.

protein-analysisproperty-prediction+3
Aggrescan3D

Aggrescan3D

Static-mode Aggrescan3D analysis for per-residue aggregation propensity from a single protein structure.

protein-analysisproperty-prediction+3
AllMetal3D

AllMetal3D

Predict metal and water binding sites in protein structures using 3D convolutional neural networks (AllMetal3D + Water3D).

structure-analysisdeep-learning+3
PROPKA 3

PROPKA 3

Predict pKa values of ionizable groups in proteins and protein-ligand complexes from 3D structure. PROPKA calculates environment-driven pKa shifts for standard ionizable residues, terminal groups, and supported ligand atom types.

protein-analysisproperty-prediction+3
SuperWater

SuperWater

Predict protein hydration sites from a structure using a diffusion model with ESM features and a confidence-filtering head.

structure-analysisai-powered+4
TNP

TNP

Profile nanobody developability with the Therapeutic Nanobody Profiler, including CDR geometry, surface hydrophobicity and charge, clinical-reference flags, and predicted structures.

protein-analysisproperty-prediction+4