# Semaglutide structure coverage: ProteinIQ audit

We analysed the deposited coordinate models in wwPDB entries **7KI0** and
**4ZGM**, retrieved from RCSB PDB on September 20, 2026. This is a targeted,
reproducible audit of two previously published structures, not new experimental
research, a survey of all semaglutide structures, or a potency comparison.

## Inputs and attribution

- 7KI0: Zhang et al., *Cell Reports* (2021), DOI
  [10.1016/j.celrep.2021.109374](https://doi.org/10.1016/j.celrep.2021.109374).
  The snapshot's latest revision is May 14, 2025.
- 4ZGM: Lau et al., *Journal of Medicinal Chemistry* (2015), DOI
  [10.1021/acs.jmedchem.5b00726](https://doi.org/10.1021/acs.jmedchem.5b00726).
  The snapshot's latest revision is January 10, 2024.

`sources.json` records download URLs and SHA-256 hashes of the uncompressed
mmCIF files. `7KI0.cif.gz` and `4ZGM.cif.gz` preserve those exact public records.
The underlying structures belong to their original investigators; the
coordinate-coverage comparison and interpretation are our analysis.

## Method

1. Select the single polymer entity whose description contains “semaglutide”.
   Both entries contain one copy. Preserve both mmCIF label and author chain IDs.
2. Read declared peptide residues from `_entity_poly_seq`. Count residues, not
   characters in a sequence string: `(AIB)` is one nonstandard residue.
3. Select coordinate model 1 and atoms with occupancy greater than zero. Exclude
   hydrogen/deuterium atoms. A residue is modeled when at least one remaining
   atom has that peptide's `label_asym_id` and `label_seq_id`. Collapse alternate
   conformations when counting residues; do not generate symmetry copies.
4. Map to author numbering with `_pdbx_poly_seq_scheme.pdb_seq_num` and compare
   missing positions with `_pdbx_unobs_or_zero_occ_residues`. The script fails
   if they disagree.
5. Inspect `_struct_conn` for explicit external covalent partners of the peptide.
   Report their chemical-component identities, formulas, attachment atoms and
   modeled heavy-atom counts. Do not assume every ligand in the file belongs to
   semaglutide, or infer a bond merely from spatial proximity.

## Findings and limits

| Entry | Label / author chain | Declared residues | With coordinates | Missing author positions | Aib coordinates |
| --- | --- | ---: | ---: | --- | --- |
| 7KI0 | E / P | 31 | 30 | Gly37 | Present |
| 4ZGM | B / B | 31 | 28 | His7, Aib8, Glu9 | Absent |

In 7KI0, `_struct_conn` links Lys26 NZ (label position 20) to component WF1
(label chain G), a linker fragment whose component formula is C12 H24 N2 O7.
Twenty positive-occupancy heavy atoms are modeled for this component. This is
not a representation of the complete C18 fatty-diacid modification. No other
external covalent partner is recorded for the peptide.

4ZGM is explicitly a semaglutide **peptide-backbone** complex. No external
covalent partner of the peptide is recorded. The separate 32M components are
not lipidation of the peptide: the file does not record a covalent linkage to
them. We do not identify semaglutide modifications from ligand presence alone.

Coverage of a residue does not establish that all its atoms are resolved.
Likewise, absence of coordinates does not establish absence from the sample.
The count does not assess density, local resolution, occupancy uncertainty,
receptor construct completeness, signaling, binding affinity or clinical effect.
The entries contain different constructs and experimental methods and should
not be ranked for overall quality using this comparison.

## Reproduce

Download this directory's `analyze.py`, `sources.json` and two `.cif.gz` files
into one directory. The original run used Python 3.13 and Biopython 1.85.

```sh
python -m pip install biopython==1.85
python analyze.py --output-dir ./reproduced
```

The script verifies the frozen input hashes before writing `results.json` and
`coverage.csv`. The latter has one row per declared peptide residue (62 rows).
No ProteinIQ hosted jobs, credits, structure repair or predicted coordinates
were used. The published chart reads the reported counts; its captions retain
the residue-level counting rule.
