# Incretin molecular comparison: inputs and method

This dataset defines a comparison of four complete incretin-drug molecules and
three reference models. The input records were retrieved on September 21, 2026.
The calculations are intended to describe molecular composition, not clinical
effectiveness, affinity, absorption, dosing or safety.

## Scope

The four drugs are semaglutide, liraglutide, tirzepatide and exenatide. We selected
these named examples, rather than sampling all incretin therapies. Each input is
one covalently connected molecule with neutral formal charge, without counterions,
water, formulation ingredients or an absorption enhancer. Neutral charge here is
a property of the input representation, not a prediction of charge in blood.

The three reference models are:

1. Semaglutide's 31-residue backbone with Aib8 and Arg34 retained, but with the
   entire Lys26 linker/lipid branch absent and the free lysine amine restored.
2. The same unlipidated backbone with alanine at position 8 instead of Aib.
3. Native GLP-1(7-37), with lysine at position 34 and alanine at position 8.

These reference models have free amino and carboxyl termini. The native model is
GLP-1(7-37), not the distinct GLP-1(7-36)amide form. The two modified unlipidated
models are computational controls, not approved medicines or experimental samples.

## Chemical identity checks

`sources.json` identifies exact PubChem CIDs, retrieval URLs and SHA-256 hashes.
The adjacent JSON snapshots preserve the retrieved SMILES, formulas and InChIKeys.
`prepare-inputs.py` checks each snapshot's hash, formula, InChIKey, connectedness
and formal charge. It reconstructs peptide backbones from their specified
sequences, introduces Aib explicitly, and verifies the lipid attachment site and
branch graph. Tirzepatide and exenatide retain their C-terminal amides.

We selected semaglutide CID 154733530 and tirzepatide CID 166567236 by comparison
with the manufacturers' structural formulas. Semaglutide CID 56843331 has a
different linker-glutamate stereochemical assignment; tirzepatide CID 156588324
has different branch connectivity. A name or elemental formula alone is therefore
insufficient to select the intended input. These checks support this small input
set; they are not a systematic assessment of database accuracy.
The two `excluded-*-pubchem.json` snapshots preserve the alternative records;
their retrieval URLs and hashes are listed separately in `sources.json`.

The exenatide input retains its downloaded neutral histidine tautomer. Its standard
InChIKey agrees with the model reconstructed from the labeled amidated sequence.
We do not enumerate physiological protonation states or tautomers.

Structural references:

- [Ozempic prescribing information, section 11](https://www.novo-pi.com/ozempic.pdf)
- [Victoza prescribing information, section 11](https://www.novo-pi.com/victoza.pdf)
- [Mounjaro original FDA label, section 11 and structural formula](https://www.accessdata.fda.gov/drugsatfda_docs/label/2022/215866s000lbl.pdf)
- [Byetta prescribing information, section 11](https://dailymed.nlm.nih.gov/dailymed/drugInfo.cfm?setid=53d03c03-ebf7-418d-88a8-533eabd2ee4f)

The archived Mounjaro label is used to verify chemical structure, not current
prescribing guidance.

## Reproduce the inputs

Download this directory's four PubChem JSON snapshots, `sources.json` and
`prepare-inputs.py` into one directory. The preparation used Python 3.13 and
RDKit 2025.03.5.

```sh
python -m pip install rdkit==2025.3.5
python prepare-inputs.py --output-dir ./prepared
```

The script writes `verified-inputs.json` and a seven-row, headerless `inputs.tsv`
with names followed by tab-separated SMILES. All molecular changes are explicit
in the code; no missing chemical modifications are inferred from a FASTA file.

## Completed ProteinIQ calculation

We executed ProteinIQ's unchanged `process_molecular_descriptors` calculation
function locally with RDKit 2025.03.5 and the **Physicochemical** descriptor set.
The seven inputs completed successfully. `tool-result.json` preserves the native
output; `execution.json` records the backend file and function hashes, repository
commit, software version, settings and input checksum. This was local backend
execution, not a hosted website job. No credits were used and no hosted job ID
was created. It does not exercise website admission, billing or result sharing.
The raw result's `executionTimeSeconds` field is supplied by the implementation;
we do not use it as a measured runtime benchmark.

We obtained 4,113.64 g/mol for complete semaglutide and 3,397.76 g/mol for the
unlipidated Aib8/Arg34 backbone. Their net difference is 715.88 g/mol, or 17.4026%
of the complete molecule's calculated mass. The net elemental increment is
C35H61N3O12, including 50 additional heavy atoms (atoms other than hydrogen).
Calculated topological polar surface area increases by 201.90 square angstroms,
with nine additional hydrogen-bond acceptors and five additional donors under
RDKit's definitions. These are graph-derived descriptors, not measurements of
solubility, exposed surface area, albumin binding or absorption.

`comparison.csv` retains seven rows. The article's four-medicine comparison uses
only the four complete drug structures; the three reference models belong to
the modification breakdown. `summary.json` supplies the mass-composition chart.

## Reproduce and verify the results

With the saved input and output files in the same directory, independently
recalculate the published descriptors and recreate the article data:

```sh
python analyze-results.py --output-dir ./reproduced
```

The verifier checks molecular identity, the input checksum and RDKit version,
then compares every published descriptor with a fresh calculation. It recreates
`comparison.csv` and `summary.json`. The two files should match the distributed
versions byte-for-byte under the recorded version.

`run-backend.py` repeats the native calculation for maintainers who have the
recorded ProteinIQ backend checkout and its dependencies. It checks both source
hashes before calling the function:

```sh
python run-backend.py --backend-path /path/to/modal_dot_com/apps/lightweight-tools/main.py --output ./reproduced/tool-result.json
```

The backend checkout is not included in this dataset. Readers can reproduce the
published descriptors with the independent verifier above or submit the inputs
through the hosted tool below.

To repeat through the hosted tool rather than the local backend:

Submit `inputs.tsv` to [Molecular Descriptors](https://proteiniq.io/app/molecular-descriptors)
with the **Physicochemical** descriptor set. Preserve the completed result export,
job identifier, settings and reported software version alongside the inputs.
Compare molecular weight, heavy-atom count, hydrogen-bond donor/acceptor counts
and topological polar surface area. Do not interpret LogP, flexibility counts or
screening heuristics as measured properties or validated clinical predictions
for these large modified peptides.

For the modification breakdown, subtract the unlipidated backbone's molecular
weight from the complete semaglutide molecule's weight using the same calculation
method. Divide that net difference by the complete molecule's weight to obtain
the percentage mass increment. This is a net covalent-modification increment,
not the molecular weight of a separately capped linker/lipid reagent. The lysine
amine loses a hydrogen upon attachment, and the chosen reference model restores it.

Average molecular weights can differ slightly from label values because of
atomic-weight conventions and rounding. Use one implementation consistently for
the comparison, retain the original exported precision, and distinguish average
molecular weight from monoisotopic mass.
