
TL;DR
- Semaglutide is a modified peptide with a 31-residue backbone. It is the active ingredient in Ozempic, Wegovy and Rybelsus.
- Two amino-acid substitutions and a lipid-containing side chain help preserve receptor activity while extending exposure; its human half-life is approximately one week.
- Our molecular calculations attribute 17.4% of complete semaglutide’s mass to the net addition of its linker and lipid, relative to the unlipidated modified backbone.
- We audited two published structures: 7KI0 contains coordinates for 30 of 31 peptide residues; 4ZGM contains 28. Neither coordinate model represents the complete lipid modification.
Semaglutide is a modified peptide that activates the GLP-1 receptor, affecting blood glucose and appetite. It is the active ingredient in Ozempic, Wegovy and Rybelsus. Its backbone contains 31 amino-acid residues, but the drug is more than a sequence: a nonstandard residue and an attached lipid-containing side chain are essential parts of its chemical identity. Those modifications help extend its half-life in humans to approximately one week. Semaglutide is not insulin; it prompts the body to release its own insulin when blood glucose is elevated.
The molecule illustrates three separate problems in drug development: activating a receptor, remaining available in the body, and reaching the bloodstream. It also highlights a practical problem for computational research, because a sequence file or a downloaded experimental structure may contain only part of the chemistry that makes semaglutide work. To show how large that gap is, we calculated how much the linker and lipid add to the complete molecule, then audited two published semaglutide structures to see which parts they actually contain.
How does semaglutide work?
Semaglutide mimics GLP-1, a hormone the gut releases after meals. By activating the GLP-1 receptor, it increases insulin secretion when blood glucose is high and reduces secretion of glucagon, a hormone that raises blood glucose. It also reduces calorie intake through effects on appetite and slows the emptying of the stomach. The receptor is the signaling target; albumin is an abundant circulating protein that affects how long the drug remains available.[1][2]
The design therefore balances two different interactions. A molecule can bind its target effectively yet disappear from the blood too quickly to be a practical medicine, so its discovery paper describes optimizing the lipid and linker for albumin binding alongside receptor potency. Target binding alone does not explain why semaglutide needs to be taken only once a week.
Prescribing information reports greater than 99% plasma-albumin binding and an elimination half-life of approximately one week. Albumin association is reversible, not a permanent chemical bond between the drug and albumin.[3][4]
Experimental structures show the other half of the mechanism, how the drug engages its receptor. Zhang and colleagues used cryo-electron microscopy, which images flash-frozen molecules, to study semaglutide bound to the GLP-1 receptor and its signaling partner, the Gs protein, including how the complex flexes. Such a structure captures receptor engagement in a defined experimental assembly; it does not show how long a dose remains in a patient.[5]
What are semaglutide's molecular weight and chemical formula?
The Ozempic label gives a molecular formula of C187H291N45O59 and molecular mass of 4,113.58 g/mol. These values describe the complete modified molecule, including the linker and lipid.[1] The rest of this section explains what that molecule is made of, starting with its amino-acid chain.
| Property | Semaglutide |
|---|---|
| Molecule type | Modified peptide; GLP-1 receptor agonist |
| Backbone length | 31 amino-acid residues |
| Chemical formula | C187H291N45O59 |
| Label molecular mass | 4,113.58 g/mol |
| U.S. brand names | Ozempic, Wegovy and Rybelsus |
What is semaglutide's amino-acid sequence?
The sequence below runs from the amino terminus to the carboxyl terminus. It uses native GLP-1 positions 7–37, so Aib8 is the second residue and Lys26 is the twentieth. Each standard letter denotes one residue; Aib denotes α-aminoisobutyric acid, and K* marks the lysine carrying the linker/lipid branch.[4]
| Native GLP-1 positions | Residues in order |
|---|---|
| 7–13 | H–[Aib]–E–G–T–F–T |
| 14–20 | S–D–V–S–S–Y–L |
| 21–27 | E–G–Q–A–A–K*–E |
| 28–34 | F–I–A–W–L–V–R |
| 35–37 | G–R–G |
The complete molecule also requires the identity and connectivity of the branch. A backbone string alone cannot specify all of semaglutide's chemistry.
How does semaglutide's structure differ from GLP-1?
Semaglutide was developed from glucagon-like peptide-1, or GLP-1, by changing two residues and modifying a lysine side chain. The discovery program sought prolonged exposure while retaining activity at the GLP-1 receptor.[6]
| Feature | Semaglutide design | Why it matters |
|---|---|---|
| Position 8 | Alanine replaced by α-aminoisobutyric acid, or Aib | Helps resist DPP-4, the enzyme that rapidly inactivates natural GLP-1 |
| Position 26 | Lysine carries a linker and C18 fatty diacid | Promotes albumin association and prolonged exposure |
| Position 34 | Lysine replaced by arginine | Leaves a single lysine site for the intended lipid attachment |
| Peptide backbone | 31 residues corresponding to GLP-1 positions 7–37 | Defines the peptide portion, not the complete modified molecule |
Aib is not one of the 20 standard amino acids that a protein sequence file, such as FASTA, can represent. The lipid is also attached to a side chain rather than added as another residue at the end of the peptide. Replacing Aib with alanine or leaving out the branch therefore changes the molecule being described, which is effectively what a standard sequence calculator does. Its mass estimate cannot be compared with the complete drug's molecular weight as though the two calculations described the same structure.
How much do the linker and lipid add?
We calculated that the linker and lipid branch adds a net 715.88 g/mol to the unlipidated peptide backbone, or 17.4% of complete semaglutide's calculated mass. Both structures in this comparison retain Aib8 and Arg34; the reference model simply restores the free amine at Lys26. The figure describes the molecule's composition, not how much the branch contributes to its clinical effect.
We used ProteinIQ's Molecular Descriptors calculator with the Physicochemical setting. The branch adds 50 heavy atoms, meaning atoms other than hydrogen, and it adds polar groups as well as the fatty chain, so describing the modification simply as a lipid leaves out part of its chemistry.
| Property | Backbone without branch | Complete semaglutide |
|---|---|---|
| Average molecular weight (g/mol) | 3,397.76 | 4,113.64 |
| Heavy atoms | 241 | 291 |
| Topological polar surface area (Ų) | 1,444.28 | 1,646.18 |
| Hydrogen-bond acceptors | 47 | 56 |
| Hydrogen-bond donors | 52 | 57 |
The net elemental increment is C35H61N3O12. It includes the linker as well as the C18 fatty diacid and accounts for the hydrogen lost from the lysine amine on attachment. The calculated complete-molecule mass, 4,113.64 g/mol, differs slightly from the label's 4,113.58 g/mol because of atomic-weight conventions and rounding; both sides of our comparison use the same implementation.
The inputs, native tool output, execution record and reproduction methods are available, including two further reference models that separate the effect of the amino-acid substitutions from that of lipidation. Descriptors like these describe structure only; they do not predict albumin binding, solubility, oral absorption or half-life.
Is semaglutide available as a pill?
Yes. U.S. semaglutide products include tablets as well as injections. Rybelsus and Ozempic tablets share a prescribing document, while Wegovy has its own tablet and injection indications and doses. Keeping a peptide in circulation for longer does not help it get there in the first place, since peptides are normally broken down in the digestive tract. The oral formulations therefore add salcaprozate sodium, known as SNAC, as an absorption enhancer.[3][2]
Buckley and colleagues investigated the formulation through clinical and preclinical studies. They found absorption in the stomach, near the tablet surface, with SNAC helping protect semaglutide through local buffering and transiently promoting passage through cells. Their results concerned the coformulated drug and enhancer, not an assumption that peptides generally pass efficiently through the digestive tract.[7]
The distinction explains why oral and injected milligram doses are not interchangeable. It also explains why knowing the molecular structure does not by itself predict the performance of a tablet: formulation and administration conditions affect delivery.[3]
How does semaglutide compare with other incretin molecules?
Semaglutide is one of several incretin medicines, drugs that mimic gut hormones released after meals. We compared its complete chemical structure with those of liraglutide, tirzepatide and exenatide using the same calculator. Across these four medicines, calculated average molecular weights range from 3,751.26 to 4,813.53 g/mol.
| Molecule | Backbone residues | Calculated mass (g/mol) | Heavy atoms | Receptor target |
|---|---|---|---|---|
| Semaglutide | 31 | 4,113.64 | 291 | GLP-1 |
| Liraglutide | 31 | 3,751.26 | 266 | GLP-1 |
| Tirzepatide | 39 | 4,813.53 | 341 | GIP and GLP-1 |
| Exenatide | 39 | 4,186.64 | 295 | GLP-1 |
Semaglutide and liraglutide have the same backbone length but different complete molecular masses. Exenatide has eight more backbone residues than semaglutide yet a similar mass. Counting amino acids alone misses the chemical modifications attached to the chain. These differences also explain why the medicines should not be described as the same molecule under different brand names.
We checked the input structures against labeled peptide sequences and attachment chemistry, rather than trusting compound names alone. The downloadable dataset includes additional calculated properties and explicitly labeled reference models. The execution record preserves the calculator version and settings. Small differences from label molecular weights reflect calculation conventions and rounding.
Where can you find a semaglutide PDB structure?
The RCSB Protein Data Bank holds experimental semaglutide structures, including 7KI0 and 4ZGM. A structure, however, contains coordinates only for the atoms the experiment could resolve. Our audit of the public mmCIF files, the standard format for these records, retrieved on September 20, 2026, found coordinates for 30 of the 31 peptide residues in 7KI0 and 28 in 4ZGM. Both records declare Aib in the sequence, but only 7KI0 contains coordinates for it.[11][12]
| Property | 7KI0 | 4ZGM |
|---|---|---|
| Experimental context | GLP-1 receptor–Gs complex; cryo-EM | Receptor extracellular domain; crystallography |
| Reported resolution | 2.50 Å | 1.80 Å |
| Peptide chain: mmCIF label / author | E / P | B / B |
| Residues declared / with coordinates | 31 / 30 | 31 / 28 |
| Missing native GLP-1 positions | Gly37 | His7, Aib8, Glu9 |
| External covalent component recorded for the peptide | WF1 linker fragment attached at Lys26 | None |
The lipid modification matters as much as the residue count. In 7KI0, we identified a recorded covalent connection between Lys26 and WF1, a linker fragment with 20 modeled heavy atoms. Its component formula contains 12 carbon atoms; it is not the full C18 fatty-diacid modification. In 4ZGM, which is explicitly a peptide-backbone complex, we found no recorded external covalent partner of the peptide. The complete lipid modification is therefore not represented in either coordinate model.
This does not make the structures defective; they capture different experimental systems and answer different questions. It does mean that a structure titled “semaglutide” should not automatically be treated as a complete model of the approved drug. Chemistry missing from a coordinate model may still have been present in the sample, only too mobile or poorly resolved to model.
How we checked the structures
We selected the semaglutide chain, compared its declared residue list with the heavy atoms actually placed in the first model (those with positive occupancy), and cross-checked missing positions against the file's missing-residue annotations. We preserved author and mmCIF chain identifiers and inspected explicit covalent connections rather than assuming every nearby ligand belonged to semaglutide.
The methods, source provenance, analysis code, residue-level CSV and full results are available. Frozen input files are included as 7KI0 and 4ZGM. We did not repair structures, predict missing atoms or run a binding calculation.
The audit covers two entries, not every semaglutide structure, and measures coordinate coverage rather than data quality or binding. The experimental structures themselves are the work of their original investigators.
What does this mean for computational studies?
The lesson from both the mass accounting and the structure audit is that the input representation should match the scientific question. A standard amino-acid sequence can support sequence-level comparisons, but it does not fully specify Aib, the linker or lipidation. A receptor-bound coordinate model can support inspection of modeled contacts, but missing chemistry remains missing unless it is explicitly reconstructed and justified.
Before modeling semaglutide, identify whether the input is the peptide backbone, a complete chemical structure, or a partial experimental model. Check nonstandard residues and covalent links, record any reconstruction, and distinguish a predicted property from an experimental measurement. The same discipline applies when comparing results from different tools.
Neither sequence-derived properties such as a stability index nor molecular descriptors calculated from a chemical structure establish semaglutide's human half-life, oral absorption or treatment benefit. Those conclusions depend on the experimental and clinical evidence discussed above.


