IPC 2.0 (isoelectric point calculator) icon

IPC 2.0 (isoelectric point calculator)

(2.0.1)

Calculate protein and peptide pI values using validated pKa scales and sequence models. Learn more

Input

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Submit job”.

What is IPC 2.0?

The isoelectric point (pI) is the pH where a protein or peptide carries zero net charge. It determines how a molecule behaves during isoelectric focusing, ion exchange chromatography, and 2D gel electrophoresis. Predicting pI from sequence requires knowing the pKa of every ionizable group, which varies depending on the pKa scale used.

IPC 2.0 (Isoelectric Point Calculator 2.0) by Kozlowski combines 19 pKa scales with machine learning to predict pI more accurately than any single Henderson-Hasselbalch calculation. The SVR models trained on 2,324 proteins and 119,092 peptides achieve RMSD values of 0.85 and 0.23 respectively. A SepConv2D deep learning model pushes peptide accuracy further (RMSD 0.22). Version 2.0 also introduced per-residue pKa prediction, estimating the dissociation constant of each ionizable residue in context using an MLP-SVR ensemble.

When to use IPC 2.0 vs simpler calculators

For a quick single-sequence estimate, the pI Calculator or Protein Parameters tool runs instantly in the browser using the Bjellqvist scale. IPC 2.0 is the better choice when:

  • Accuracy matters (SVR and DL models outperform any single pKa scale)
  • Comparing predictions across multiple scales helps assess confidence
  • Per-residue pKa values are needed for understanding charge distribution
  • Batch processing hundreds of sequences at once

How to use IPC 2.0 online

IPC 2.0 on ProteinIQ predicts the isoelectric point of proteins and peptides from amino acid sequence. Paste sequences in FASTA format (or fetch them from UniProt), choose a prediction method, and get pI values across up to 19 pKa scales plus machine learning predictions, along with per-residue pKa for every ionizable site.

Inputs

InputDescription
Protein/Peptide SequencesOne or more amino acid sequences in FASTA format. Also accepts .fasta, .fa, .fas, or .txt files. UniProt IDs can be fetched directly using the batch fetcher.

Settings

SettingDescription
PredictorPrediction method. Default: IPC1_ALL. See predictor options below.

Predictor options

PredictorBest forWhat it does
IPC1_ALLGeneral useHenderson-Hasselbalch with all 19 pKa scales (18 literature scales + ProMoST). Fastest option.
IPC2_SVR_proteinProteinsSVR trained on 2,324 proteins from SWISS-2DPAGE and PIP-DB. Reports all 19 scales plus the SVR prediction.
IPC2_SVR_peptidePeptidesSVR trained on 119,092 peptides from HiRIEF experiments. Reports all 19 scales plus the SVR prediction.
IPC2_DL_peptidePeptides ≤60 aaSepConv2D deep learning model. Highest accuracy for short peptides (RMSD 0.22).
ALLMethod comparisonRuns every predictor and all pKa scales. Useful for assessing prediction confidence.

Results

The Results tab shows a spreadsheet with one row per sequence:

ColumnDescription
IDSequence identifier from the FASTA header.
LengthNumber of amino acid residues.
MWMolecular weight in Daltons.
pI_[scale]Predicted pI for each pKa scale (e.g., pI_IPC2_protein, pI_Bjellqvist, pI_ProMoST).
pI_SVR_proteinSVR prediction optimized for proteins (when SVR or ALL method selected).
pI_SVR_peptideSVR prediction optimized for peptides (when SVR or ALL method selected).
pI_DL_peptideDeep learning prediction (when DL or ALL method selected).
Avg_pIAverage pI across all calculated scales (excluding Patrickios).

The Per-Residue pKa tab shows predicted pKa values for every ionizable site:

ColumnDescription
SequenceWhich input sequence this residue belongs to.
PositionResidue position (1-indexed), or N-term/C-term for terminal groups.
ResidueAmino acid type (D, E, H, K, Y, or the terminal residue).
pKaPredicted dissociation constant from the MLP-SVR ensemble.

Interpreting pI values

Most proteins have pI between 4 and 10. Knowing where a protein falls relative to pH 7 determines its behavior in common buffers:

pI rangeCharacterAt pH 7.4 (physiological)
<5Strongly acidicNet negative charge
5 to 7Weakly acidicSlight negative charge
7 to 9Weakly basicSlight positive charge
>9Strongly basicNet positive charge

When multiple pKa scales agree within 0.3 pH units, the prediction is likely reliable. Disagreement of more than 1 pH unit suggests the protein has unusual charge properties (e.g., many histidines) where scale choice matters. In those cases, the SVR prediction is more trustworthy because it was trained on experimental data and can capture non-additive effects.

Interpreting per-residue pKa

Per-residue pKa values show how sequence context shifts each ionizable group's dissociation constant away from its "standard" textbook value. For example, an aspartate surrounded by other negative residues will have a higher pKa (harder to deprotonate), while one near positive charges will have a lower pKa.

These values are useful for identifying residues with unusual protonation behavior, which matters for enzyme active sites, pH-dependent conformational changes, and designing mutations that shift pI.

How does IPC 2.0 work?

Henderson-Hasselbalch method

The classical approach calculates net protein charge as a function of pH using:

pH=pKa+log⁡[A−][HA]\text{pH} = \text{p}K_a + \log\frac{[\text{A}^-]}{[\text{HA}]}pH=pKa​+log[HA][A−]​

Starting at pH 6.51, a bisection algorithm adjusts pH up or down until the sum of positive charges (Lys, Arg, His, N-terminus) equals negative charges (Asp, Glu, Cys, Tyr, C-terminus) within a precision of 0.01 pH units.

Each pKa scale assigns different dissociation constants to these groups, producing different pI estimates. The Bjellqvist scale additionally uses position-specific pKa values for N-terminal and C-terminal residues (e.g., N-terminal Ala has pKa 7.59 instead of the generic 7.5), which matters for proteins starting with A, M, S, P, T, V, or E.

ProMoST uses a separate model with three pKa values per residue (N-terminal, middle, C-terminal positions), giving it slightly different behavior from the other 18 scales.

pKa scales

IPC 2.0 includes 19 pKa scales:

ScaleOrigin
IPC2_proteinOptimized for proteins (Kozlowski 2021)
IPC2_peptideOptimized for peptides (Kozlowski 2021)
IPC_proteinOriginal IPC scale (Kozlowski 2016)
IPC_peptideOriginal IPC scale (Kozlowski 2016)
ProMoSTPosition-specific model (Halligan 2004)
Bjellqvist2D electrophoresis standard with terminal corrections
EMBOSSEMBOSS software suite
DTASelectProteomics analysis
GrimsleyExperimental NMR measurements
LehningerBiochemistry textbook
SolomonProtein chemistry
SilleroTheoretical calculations
RodwellBiochemistry reference
ThurlkillNMR measurements in pentapeptides
ToselandStatistical analysis of PDB structures
NozakiModel compound measurements
DawsonData compilation
WikipediaGeneral reference values
PatrickiosSimplified (only D/E/K/R, ignores C/H/Y)

The Avg_pI column averages all scales except Patrickios, whose simplified model (ignoring cysteine, histidine, and tyrosine) makes it an outlier.

Support vector regression

The SVR models take a 19-dimensional feature vector (pI predictions from all 18 H-H scales + ProMoST) and feed it into a support vector machine with RBF kernel. The protein SVR was trained on 2,324 proteins from SWISS-2DPAGE and PIP-DB. The peptide SVR was trained on 119,092 peptides from high-resolution isoelectric focusing (HiRIEF) experiments.

Because the SVR sees all 19 scale predictions simultaneously, it learns which scales are most informative for different sequence compositions. It consistently outperforms any individual scale.

SepConv2D deep learning

The deep learning model uses a separable convolutional architecture with four input channels:

  • One-hot encoding: 22 x 60 matrix of amino acid identity
  • AAindex features: 15 physicochemical indices per position plus hydrophobicity
  • Composition vector: Amino acid counts repeated across sequence length
  • IPC predictions: pI from all scales + SVR, giving the model calibrated priors

Sequences longer than 60 residues undergo truncation that preferentially removes non-charged residues (Ala, Gly, Leu, etc.) while preserving ionizable residues that determine pI. This truncation uses random selection among non-charged positions, so predictions for long sequences may vary slightly between runs.

Per-residue pKa (MLP-SVR ensemble)

The per-residue pKa predictor estimates the dissociation constant of each ionizable residue (D, E, H, K, Y) plus the N- and C-terminal groups. It uses a stacking ensemble of 9 multilayer perceptrons (MLPs) feeding into a final SVR:

  • 6 sequence-only MLPs operating on 5-mers through 15-mers around the target residue
  • 3 AAindex-augmented MLPs operating on 3-mers through 7-mers

Each MLP encodes the local sequence context around an ionizable residue using one-hot encoding (and optionally AAindex physicochemical features). The 9 MLP predictions are stacked as features for the final SVR, which outputs a single pKa value. Charged N- and C-terminal residues are flanked with alanine padding before prediction to ensure they are treated as terminal groups.

Table of contents

Related tools

FindPept

FindPept

Match experimental peptide masses against theoretical digest fragments of a protein sequence. Identify peptides from mass spectrometry data by peptide mass fingerprinting.

protein-analysisphysicochemical-properties+2
Peptide cutter

Peptide cutter

Predict protease and chemical cleavage sites across a protein sequence for up to 39 enzymes simultaneously. Identify where each enzyme cuts, the cleavage residue, and context window around each site.

protein-analysisphysicochemical-properties+2
Aggrescan3D

Aggrescan3D

Static-mode Aggrescan3D analysis for per-residue aggregation propensity from a single protein structure.

protein-analysisproperty-prediction+3
Protein charge plot

Protein charge plot

Plot net charge vs pH for protein sequences. Visualize how protein charge changes across pH 0-14 and identify the isoelectric point (pI) where the net charge crosses zero.

protein-analysisphysicochemical-properties+2
Hydropathy plot

Hydropathy plot

Generate Kyte-Doolittle hydropathy plots to visualize hydrophobic and hydrophilic regions along protein sequences. Identify transmembrane domains and surface-exposed regions.

protein-analysisphysicochemical-properties+2
Hydrophobicity plot

Hydrophobicity plot

Generate hydrophobicity plots using 24 different amino acid scales. Visualize hydrophobic and hydrophilic regions for protein analysis, epitope prediction, and membrane protein studies.

protein-analysisphysicochemical-properties+2
Peptide mass calculator

Peptide mass calculator

Cleave a protein sequence with a chosen protease and compute the masses of the resulting peptides. Supports multiple enzymes, missed cleavages, chemical modifications, and different ion types for mass spectrometry experiment planning.

protein-analysisphysicochemical-properties+1
PROPKA 3

PROPKA 3

Predict pKa values of ionizable groups in proteins and protein-ligand complexes from 3D structure. PROPKA calculates environment-driven pKa shifts for standard ionizable residues, terminal groups, and supported ligand atom types.

protein-analysisproperty-prediction+3
Protein parameters

Protein parameters

Calculate sequence-derived protein properties including molecular weight, theoretical pI, extinction coefficients, aromaticity, secondary structure fractions, composition classes, instability, aliphatic index, and GRAVY.

protein-analysisphysicochemical-properties+1
Protein scale profiler

Protein scale profiler

Generate amino acid property profiles using 42 different scales spanning hydrophobicity, secondary structure propensity, flexibility, polarity, surface accessibility, antigenicity, and more.

protein-analysisphysicochemical-properties+2