7.3: Glycoconjugates - Proteoglycans, Glycoproteins, Glycolipids and Cell Walls
- Page ID
- 14957
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)
(Learning goals written by Claude, Sonnet 4.6, Anthropic)
Many proteins, especially those destined for secretion or insertion into membranes, are post-translationally modified by the attachment of carbohydrates. They are usually attached through either Asn or Ser side chains. Carbohydrate modifications on the protein appear to be involved in recognizing other binding molecules, preventing aggregation during protein folding, protecting from proteolysis, and increasing the protein's half-life. In contrast to a protein sequence determined by a DNA template, sugars are attached to proteins by enzymes that recognize appropriate sites on proteins and attach the sugars. Since there are many sugars with many functional groups that can serve as potential attachment sites, the structures of the oligosaccharides attached to proteins are enormously varied, complex, and hence "information-rich" compared to linear or folded polymers like DNA and proteins.
N-linked Glycoproteins
These contain carbohydrates attached through either a GlcNAc or GalNAc to an Asn in an X-Asn-X-Thr sequence of the protein. There are three types of N-linked glycoproteins: high mannose, complex, and hybrid. They all contain the same core oligosaccharide - (Man)3(GlcNAc)2 attached to Asn as shown in Figure \(\PageIndex{1}\).
Table \(\PageIndex{1}\) below shows the SNFG representation for the main core and variant glycans of N-linked glycoproteins. Note that the designation of α2 implies an α(1→2) linkage. Unless otherwise stated, the linkage is presumed to start from carbon 1.
| Core | |
| High mannose | |
| Complex | |
| Mixed hybrid |
Table \(\PageIndex{1}\): SNFG representation for the main core and variant glycans in N-linked glycoproteins
Complex N-linked glycans don't contain mannose outside the core glycan and have GlcNAc attached to the branching mannoses in the core structure. The complex glycan shown above has a Gal(β1,4)GlcNAc sequence, which could be named the disaccharide lactosamine. Often, lactosamine repeats in the sequence.
Hybrid glycans have both unsubstituted terminal mannoses (as in the high-mannose type) and substituted mannoses with an N-acetylglucosamine attached (as in the complex type. GlcNAc residues added to the core in the hybrid and complex N-glycoproteins are called antennae. Figure \(\PageIndex{2}\) shows an example of a biantennary N-linked glycan with two GlcNAc branches linked to the core. The core is outlined in red, and the two GlcNAcs are labeled 1 and 2.
Complex glycans also have bi-, tri-, and tetraantennary forms and comprise most N-glycans. As shown in Table \(\PageIndex{1}\) above, complex N-linked glycans usually end with sialic acid residues. About 50% of the surface area of the COVID-19 SARS-CoV-2 spike protein is covered with glycans, as shown in the model structure in Figure \(\PageIndex{3}\). The protein surface is gray, and the glycans (biantennary LacNAc N-glycans) in spacefill CPK with carbon in cyan.
In the hybrid oligosaccharide shown above, one terminus contains Gal(β1,4)GlcNAc. However, in all other mammals except man, apes, and Old World monkeys, an additional Gal is often connected in an α1,3 link to the Gal to give a terminus of Gal(α1,3)Gal(β1,4)GlcNAc. These animals have an additional enzyme, an α1,3-galactosyltransferase. Bacteria also have this enzyme, and since we have been exposed to this link through bacterial infection, we mount an immune response against it. Why is this important? Pig hearts are similar to human hearts, so they might be good candidates for human transplantation (xenotransplants). However, the Gal-α1,3-Gal link is recognized as foreign, and we mount a significant immune response against it. Several biotech firms are trying to delete the pig α1,3 Gal transferase, which would prevent the addition of the terminal Gal and make them good donors for transplanted hearts.
Figure \(\PageIndex{4}\) shows an interactive iCn3D model of an N-linked glycoprotein, human beta-2-glycoprotein-I (Apolipoprotein-H) (1C1Z).
The glycan structures for the beta-2-glycoprotein-I are shown below. Identify the monosaccharides in each and specify to which asparagine they are linked.
- Answer
-
Here is an interactive iCn3D image showing the protein and attached glycans using SNFG notation.
Figure \(\PageIndex{5}\) shows an interactive iCn3D model of the GP120 HIV protein that contains high mannose, complex, and hybrid N-linked glycans. Most glycoproteins in the Protein Data Bank do not contain attached glycans. The glycans here were added using the program GlyProt at 3 of 17 possible Asn residues that would presumably have attached glycans. Use your mouse or keypad to hover over the monomers in the attached glycans. Abbreviations for the given residues in the model are: adm = alpha-D-Man, bdg = beta-D-Glc or Gal, adn = alpha-D-neuraminidase.
The coronavirus pandemic has been deadly (over 1.1 million deaths in the USA alone and 19-36 million around the world). However, the 1918 influenza pandemic was far worse per capita, with an estimated 650,000 deaths in the USA and 50 million around the world in a population less than 1/3 of the present. An additional 3 million deaths in the USA were probably prevented by vaccination as of December 2022. Many deaths in the developing world would have been prevented if wealthy nations had allocated more resources to vaccine production and distribution. The virus's evolution in unvaccinated areas might come back to haunt wealthy countries if current vaccines become ineffective against mutants. A worse pandemic might await us. An avian version of the influenza virus (H5N1), presently endemic in wild birds and now found in mink populations, has infected 240 people as of January 2023 and killed 56% of them. A quick note: the 1918 pandemic affected youth the most.
Figure \(\PageIndex{6}\) shows the simple yet deadly influenza virus. It interacts with human cells through a surface protein, hemagglutinin (HA).
Figure \(\PageIndex{6}\): Alpha Influenza virus. https://viralzone.expasy.org/6. Creative Commons Attribution 4.0 International (CC BY 4.0) License.
Note the similarities and differences with the SARS-CoV-2 virus, shown below.
SARS-Covid-2 virus. https://viralzone.expasy.org/764. Creative Commons Attribution 4.0 International (CC BY 4.0) License.
The virus binds to host cells through the interaction of HA with cell surface carbohydrates. Once bound, the virus internalizes, ultimately releasing the viral RNA genome into the host cell.
The hemagglutinin protein is the most abundant protein on the viral surface. 15 avian and mammalian variants have been identified (based on antibody studies). Only three have adapted to humans in the last 100 years, giving pandemic strains H1 (1918), H2 (957), and H3 (1968). Three recent avian variants (H5, H7, and H9) have jumped directly to humans recently, but have low human-to-human transmissibility.
The influenza hemagglutinin protein has the following characteristics:
- the mature form is a homotrimer (3 identical protein subunits), MW 220,000, with multiple sites for covalent attachment of sugars. Hemagglutinin is a glycoprotein.
- each monomer is synthesized as a single polypeptide chain precursor (HA0), which is cleaved into HA1 and HA2 subunits by the protease trypsin in epithelial cells of the lung.
- structure known for human (H3), swine (H9), avian (H5) subtypes.
Figure \(\PageIndex{7}\) below shows an interactive iCn3D model of an influenza hemagglutinin trimer (6HJQ). The white backbone traces on the membrane's extracellular (red) side are antibody molecules used to stabilize the structure for crystallization.
Figure \(\PageIndex{7}\): Influenza Hemagglutinin Trimer (6HJQ). (Copyright; author via source). Click the image for a popup or use this external link: https://structure.ncbi.nlm.nih.gov/i...EhePsF4BiEjwA8
The protein is a trimer of a heterodimer, (HA1:HA2)3. Each heterodimer consists of two distinct chains, HA1 and HA2. Together they form a globular head and stalk region, the latter crossing the membrane. HA1 primarily interacts with sialic acid on host cell surface proteins, while HA2 is more involved in viral fusion with host cells and internalization.
Hemagglutinin binds to sialic acid (Sia) covalently attached to many cell membrane glycoproteins. The sialic acid is usually connected through an α(2,3) or α(2,6) link to galactose on N-linked glycoproteins. The subtypes found in avian (and equine) influenza isolates bind preferentially to Sia (α2,3) Gal, which predominates in the avian GI tract, where viruses replicate. Human influenza isolates prefer Sia α(2,6)
Sia (α2,3) Gal predominates in the avian GI tract, where viruses replicate. Human influenza isolates preferentially bind Sia α(2,6)Gal. The human H1, H2, and H3 viruses (causes of the 1918, 1957, and 1968 pandemics) recognize Sia α(2,6)Gal, the major form in the human respiratory tract. The swine influenza HA binds to Sia α(2,6) Gal and some Sia (α2,3) Gal, both found in swine. The structures of the Sia-Gal disaccharide are shown in Table \(\PageIndex{2}\) below.
| Sia α(2,6) Gal (Human) | Sia α(2,3) Gal (Avian and some Swine) |
|
|
|
|
(made with Sweet, with an OH, not AcNH on sialic acid on C5) |
(made with Sweet, with an OH, not AcNH on sialic acid on C5) |
Table \(\PageIndex{2}\): Structures of Sia α(2,6) Gal (human) and Sia α(2,3) Gal (avian/swine)
The H5N1 avian flu H5N1) virus is deadly but presently lacks human-to-human transmissibility. Why? One reason is that it appears to bind deeply in the lungs and is not easily released by coughing or sneezing. It appears that cell surface glycoproteins deeper in the respiratory tract have Sia (α2,3) Gal which accounts for this pathology.
Before it leaves the cell, the virus forms a bud on the intracellular side of the host cell, with the HA and NA in the host cell membrane. The virus in this state would not leave the cell, as its HA molecules would interact with sialic acid residues in the host cell membrane, thereby holding it in place. Neuraminidase hydrolyzes sialic acid from cell-surface glycoproteins, allowing the virus to complete budding and be released from the cell as new viruses. The drugs Oseltamivir (Tamiflu) and Zanamivir (Relenza) bind to and inhibit neuraminidase, an enzyme required for viral release from infected cells. Tamiflu appears to be active against N1 of the current H5N1 avian influenza viruses. Governments across the world are hopefully stockpiling this drug in case of a pandemic caused by the avian virus jumping directly to humans and becoming transmissible from human to human.
Recent Updates: 6/26/26
3D N-linked Glycan Structures
As we mentioned before, a few 3D conformation structures of glycans are known, as they are so complex, branched, chemically modified, flexible, and not defined by a "genetic" template but rather by synthetic enzymes. A high-resolution structure of the native T-mastigoneme central tubes from O. danica was determined using cryo-EM and modeling. It is formed from six OChromonas Mastigoneme-related glycoproteins (OCM1 to OCM6). OCM5 and OCM6 form helical tubes with glycans covering their surface. The sixe proteins build the central tube of the tubular mastigoneme in the golden alga Ochromonas danica. Three new N-linked glycans were determined that likely exist in near-identical conformations in other proteins. Each of the glycans described below has a distinct structural role.
N-linked Glycostitches
This glycan is not covalently attached to the canonical NXS/T as described above. Rather, they are attached to an Asn in an Ala-Asn-Asp (AND) sequence. Instead of a Man3GlcNAc2 core, they are attached by a glycosidic bond to D-xylose followed by short linear α-1,2-linked D-Man (3mannoses in OCM3, 6in OCM4). This structure forms a short helix that "stitches" together neighboring protein domains, lining the interface between the two halves of a repeat unit and mediating intrarepeat OCM–OCM contacts via direct and water-mediated H-bonds. PDB examples: N70-glycan in OCM4. The distal end interacts noncovalently via hydrogen bonding with a neighboring subunit. The glycostitch on OCM3 interacts with OCM1 and OCM2. The OCM4 stitch interacts with OCM5 and OCM6.
Figure \(\PageIndex{x1}\) shows an interactive iCn3D model of an N-glycan glycostitch from tubular mastigoneme (the central tube) from golden algae (9UFE) The
Figure \(\PageIndex{x1}\): N-glycan glycostitch from tubular mastigoneme protein (the central tube) from golden algae (9UFE). (Copyright; author via source). Click the image for a popup or use this external link: https://www.ncbi.nlm.nih.gov/Structure/icn3d/share2.html?f8bb39d6b44814302a168018813e5085 (Slow Load)
The subunits are shown in a cartoon with different colors.
N-linked Glycobridges
This high mannose and branched structure has an Man3GlcNAc2 core. Its function appears to be stabilize the protein within and across repeat domains in OCM6. The 3D structure of the N240-glycan in OCM6 is almost identical to the N332 glycan of HIV-1 gp120 protein (RMSD 0.93 Å). Figure \(\PageIndex{x2}\) shows an interactive iCn3D model of an N-glycan glycobridge from glycosylated engineered gp120 outer domain (3TYG).
Figure \(\PageIndex{x2}\): N-glycan glycobridge from glycosylated engineered gp120 outer domain (3TYG). (Copyright; author via source). Click the image for a popup or use this external link: https://www.ncbi.nlm.nih.gov/Structu...44b16e32eaa3c3
The magenta chain is the gp120, and the brown and blue chains are neutralizing antibody chains. The N332 in the wild-type GP120 corresponds to N121 in the engineered structure above.
N-linked Glyco-islets
This more-inferential motif has a Man₃GlcNAc₂ core linked at the canonical NXS/T site, but also includes a D-xylose and a conserved 6-phospho-D-mannose (6-Pi-Man) that, along with other amino acid side chains, participates in binding Ca2+. The term "islets" refers to their clustering on the protein surface. This logically increases the surface's polar character and likely confers resistance to proteases. No similar 3D structure is found in the PDB, possibly because they would resolve only in a protein complex. Figure \(\PageIndex{x3}\) shows an interactive iCn3D model of an N-glycan glyco-islet from OCM5 in tubular mastigoneme from golden algae (9UFE).
.png?revision=1&size=bestfit&width=430&height=389)
Figure \(\PageIndex{x3}\): N-glycan Glycoislet from OCM5 in tubular mastigoneme (the central tube) from golden algae (9UFE). (Copyright; author via source). Click the image for a popup or use this external link: https://www.ncbi.nlm.nih.gov/Structure/icn3d/share2.html?b8689e3feae653c3308937931f9702f9
O-linked Glycoproteins
The CHOs are usually attached from a Gal (β 1,3) GalNAc to a Ser or Thr of a protein, as shown in Figure \(\PageIndex{8}\).
Figure: O-linked Glycoproteins
The blood group antigens (CHO groups on cell-surface proteins or lipids) are examples. The sugars shown as chairs (in contrast to structures found in many texts) in Figure \(\PageIndex{9}\) are the blood group antigens. They are attached to a core heterosaccharide (a red ellipse below), which is connected to either a membrane glycoprotein or glycolipid.
Figure \(\PageIndex{10}\) shows the SNFG representation for the A antigen in the glycolipid form.
The trimeric branched residues on the left-hand side represent the A antigen shown above. The red triangle is L-fucose. Yellow represents galactose (GalNac), while blue represents glucose (GlcNAc).
Proteoglycans
Some proteins are so modified with CHOs that they contain more CHOs than amino acids. Proteins linked to glycosaminoglycans are called proteoglycans (PGs). They consist of a core protein linked to one or more glycosaminoglycans. GAGs are linear sulfated glycans, which we described earlier. The structures of a few proteoglycans are known. The GAGs are O-linked to the protein, typically to a Ser of a Ser-Gly dipeptide, often repeated in the protein. Some of the proteoglycans also contained N-linked oligosaccharide groups. Figure \(\PageIndex{10}\) represents a proteoglycan structure.
PGs can be soluble and found in the extracellular matrix, or they can be integral membrane proteins. There are about 43 genes for proteoglycans. Differential splicing of the RNA transcripts gives rise to soluble and transmembrane forms. Given the diversity of sugars and varying degrees of sulfation, the CHO portion of PGs provides an incredible array of binding structures at or near the cell surface. Figure \(\PageIndex{11}\) shows the variety of proteoglycans found in mammalian cells. PGs help form the extracellular matrix, which provides a rich binding environment between cells.
One PG, syndecan, binds through its intracellular domain to the cell's cytoskeleton while interacting with another protein, fibronectin, in the extracellular matrix. Fibronectin also binds other molecules, regulating cellular growth and other interactions. PGs act as glue, connecting the cell's extracellular and intracellular functions. There are four different core syndecan proteins (SDCs 1–4), with SDC4 lacking the cytoplasmic and transmembrane domains, so it is a soluble form in the intracellular matrix. The glycan components of syndecans are mostly heparan sulfate, while SDC1 and SDC3 also have two chondroitin sulfate chains.
Most proteins bind PGs through a PG binding motif of BBXB or BBBXXB, where B is a basic amino acid. Some proteins bind to specific sequences in specific GAGs. For instance, antithrombin 3, an inhibitor of blood clotting, binds specifically to heparin. This enhances its interaction with clotting proteins such as thrombin and Factor Xa. Figure \(\PageIndex{12}\) shows an interactive iCn3D model of a five-residue fragment of heparin interacting with the key amino acids' side chains of Factor Xa (2gd4).
The extracellular matrix (ECM) might appear to be a nondescript mess for those more chemically oriented, since chemists are used to well-defined structures. Figure \(\PageIndex{13}\) shows a cartoon of the ECM and may clarify the components. Few structure files exist for them, given the inherent flexibility of the glycan components.
Cell Walls and Glycolipids
In contrast to eukaryotic cells, bacteria, and plant cells have a cell wall in addition to a lipid bilayer membrane. These are essentially carbohydrate polymers that determine cell shape, affording protection against external pathogens, hypotonic conditions, and high internal osmotic pressures, thereby preventing cell swelling and bursting. This is especially important in plants, which need strength and rigidity against the "turgor" pressure of the aqueous cytoplasm against the cell membrane. This prevents wilting in plants. The cell walls of plants and, probably, bacteria are involved in cell signaling across the cell membrane.
Bacterial Cell Walls
Two types of cell walls occur.
a. Gram-positive bacteria-
These bacteria can be stained with a Gram stain. The wall consists of a GlcNAc (β 1,4) MurNAc repeat. (GlcNAc is often abbreviated as NAG, while MurNAc is abbreviated as NAM.) This is similar to the GlcNAc (β 1,4) GlcNAc homopolymer chitin, except that every other GlcNAc contains a lactate molecule covalently attached in an ether-linkage to the C3 hydroxyl to form the monomer N-Acetylmuramic acid. A pentapeptide (Ala-D-isoGlu-Lys-D-Ala-D-Ala) is attached through an amide link to the carboxyl group of the lactate in MurNAc. A pentaglycine bridge covalently connects the GlcNAc (β 1,4) MurNAc strands through the epsilon amino group of the pentapeptide Lys on one strand and the terminal D-Ala of a pentapeptide on another strand. A small part of the structure of a gram-positive bacterial cell wall is shown in Figure \(\PageIndex{14}\). It shows one repeating GlcNAc-MurNAc disaccharide unit in front (darker) and one in the back (lighter) connected through the peptides shown.
The SNFG representation of a larger section of the gram-positive cell wall is shown in Figure \(\PageIndex{15}\).
One final structure is found in Gram-positive peptidoglycan cell walls. Teichoic acids are often attached to carbon 6 of MurNAc. Teichoic acid is a polymer of glycerol or ribitol with alternative GlcNAc and D-Ala linked to the middle C of the glycerol. Multiple glycerols are linked through phosphodiester bonds. These teichoic acids often account for 50% of the cell wall's dry weight and present a foreign (or antigenic) surface to infected hosts. These often serve as receptors for viruses that infect bacteria (called bacteriophages). Its structure is illustrated in Figure \(\PageIndex{16}\).
Notice that all monomeric units of peptidoglycan and attached teichoic acid derivatives are covalently attached to form one large molecule comprising the entire cell wall! This structure, along with the Gram-negative cell wall structures, is the largest single macromolecule in nature.
b. Gram-negative bacteria
These bacteria can NOT be stained with a Gram stain. The wall consists of the same structure as in Gram-positive bacteria. However, the GlcNAc (β 1,4) MurNAc strands are covalently connected through a direct amide bond between a derivative of Lys, meso-diaminopimelic acid (m-A2pm), on one peptide strand and to the last D-Ala of a pentapeptide on another strand. (i.e., there is no penta-Gly spacer). The connector peptide is Ala-D-isoGlu-m-A2pm-D-Ala-D-Ala
m-A2pm replaces Lys 3 of the peptide in most Gram-negative species and Gram-positive bacteria of the genus Bacillus and mycobacteria. The stereochemistries at each chiral center are different (R and S), but because the molecule has a plane of symmetry, it is an example of a meso-compound, a diastereoisomer of a molecule that lacks an enantiomeric form. The structure is shown in Figure \(\PageIndex{17}\).
A small part of the structure of a Gram-negative bacterial cell wall is shown in Figure \(\PageIndex{18}\).
Figure \(\PageIndex{19}\) shows an image of a computed model (not a crystal or NMR structure) of the Gram-negative peptidoglycan of E. Coli. Jame Gumbart kindly provided the PDB coordinates. The repeating (GlcNAc-MurNAc)n are shown in sticks, alanines in red spheres, m-A2-pm (meso-diaminopimelic acid) in gray spheres, and DiG in orange spheres. The glycine pentapeptide is not shown since it is not found in Gram-negative bacteria.

Figure \(\PageIndex{19}\): Part of a Gram-negative peptidoglycan of E. Coli. PDB coordinates kindly provided by Jame Gumbart.
Follow these instructions to get an interactive iCn3D model of the structure.
- open iCn3D
- Download this file to your computer's download folder. IMPORTANT: If the file opens as an image in a new browser window, right-click the image and save the file to download it!
- File, Open File, iCn3D png (appendable), and choose the downloaded file (it takes a while).
In addition, Gram-negative bacteria don't have teichoic acid polymers. Rather, they have a second, outer lipid bilayer. The cell wall peptidoglycan (PG) is sandwiched between the inner and outer bilayers. The space between the lipid bilayers is called the periplasm. The outer leaflet of the outer membrane is coated with a lipopolysaccharide (LPS) (a glycolipid) of varying composition. The LPS determines the antigenicity of the bacteria. The different LPS are called the O-antigens. Figure \(\PageIndex{20}\) shows the structure of the Gram-negative bacterial membrane organization. (In the figure, PS is LPS, PG is peptidoglycan). The LPS in the outer leaflet is amphiphilic, with nonpolar acyl chains forming a more classic bilayer, and the inner leaflet phospholipid and polar/charged sugars form the LPS fringe. The extra membrane of Gram-negative bacteria makes them the major source of antibiotic resistance.
A detailed view of the structure of the lipopolysaccharide (LPS) from Salmonella typhimurium is shown in Figure \(\PageIndex{21}\) below.
Figure \(\PageIndex{21}\): Lipopolysaccharide (LPS) from Salmonella Typhimurium
Recent Updates: 11/8/24
LPS is synthesized in the inner leaflet of the inner membrane and must translocate all the way to the outer leaflet of the outer membrane. Movement requires an ATP-binding cassette transporter LptB2FG. Figure \(\PageIndex{22}\) shows the machinery used to move LPS to the outer membrane. We'll discuss membrane proteins and transport thoroughly in Chapter 11.
Figure \(\PageIndex{22}\): LPS transport from the IM to the OM by the trans-envelope complex LptABCDEFG. Dong, H., Zhang, Z., Tang, X. et al. Structural and functional insights into the lipopolysaccharide ABC transporter LptB2FG. Nat Commun 8, 222 (2017). https://doi-org.ezproxy.csbsju.edu/1...67-017-00273-5. Creative Commons Attribution 4.0 International License. http://creativecommons.org/licenses/by/4.0/.
Legend: LPS is extracted from the periplasmic side of the IM by the ABC transporter LptB2FG, and is delivered to an IM protein LptC, which forms a complex with LptB2FG. LptC comprises a single membrane-spanning domain and a large periplasmic domain, forming a periplasmic bridge with LptA and the N-terminal domain of LptD. LPS is then inserted into the OM by the LptD/E complex. LPS contains O-antigen, core oligosaccharide, and lipid A components, of which the O-antigen has 4-40 O-antigen repeat units. Ra-LPS, rough LPS, Kdo, 3-deoxy-D-manno-oct-2-ulosonic acid, Hep, L-glycero-D-manno-heptose, Glc, D-glucose, Gal, D-galactose.
Figure \(\PageIndex{23}\) shows an interactive iCn3D model of the E. Coli lipopolysaccharide ABC transporter LptB2FG (6MHU). The left panel shows the membrane protein that binds and transports the LPS (spacefill) from the inner to the outer membrane. The right panel shows the bound LPS. Just part of the inner core is shown since other parts were likely too flexible to be observed. The outer O antigen is not shown since a mutation in the protein prevented it. Six acyl chains are evident. The phosphorylated GlcNAc, KDO, and HEP "layers" are labeled in the right panel.
|
Click the image for a popup or use this external link: https://structure.ncbi.nlm.nih.gov/i...CoPtdH2RVY82P6. Click Style, Background, Transparent in the iCn3D window for a better view. |
Click the image for a popup (instructions below) |
Figure \(\PageIndex{23}\): E. Coli lipopolysaccharide ABC transporter LptB2FG (6MHU) (left panel) and the bound LPS (right panel). (Copyright; author via source).
To see an interactive iCn3D model of the E. Coli LPS from the right panel, follow these steps:
- open iCn3D
- Download this file to your computer's download folder. IMPORTANT: If the file opens as an image in a new browser window, right-click the image and save the file to download it!
- File, Open File, iCn3D png (appendable), and choose the downloaded file.
c. Archaeal Cell Membranes and Walls
We have already discussed that the lipids in Archaeal cell membranes contain L (instead of D) glycerol derivatives and that ether links (more stable in reactive environments) replace ester links with isoprenoid (sometimes branched) chains, replacing fatty acid chains. The cell wall is also quite different, and some don't have one. The type of cell wall depends on the environmental need for stability. They don't contain peptidoglycans. Figure \(\PageIndex{24}\) shows four different types.
Some differences include the presence of
- pseudomurein - This is the closest to the peptidoglycans presented above. Instead of repeating disaccharide units of (NAM-NAG)n, they have a repeating disaccharide unit of N-acetylalosaminuronic acid (NAT)-NAG. The structure of NAT is shown in Figure \(\PageIndex{25}\).
Figure \(\PageIndex{25}\): N-acetylalosaminuronic acid (NAT)compared to N-acetylmuramic acid (NAM)
- methanochondroitin - This is similar to the glycosaminoglycan chondroitin sulfate
- S-Layer
- Sheath/S-Layer
d. Plant Cell Wall
If you thought bacterial cell walls were complicated, wait until you see plant cell walls! There are about 35 different types of plant cells, and each may have a different cell wall depending on the local needs of a given cell. Cells synthesize thin cell walls that remain thin as the cell grows.
Figure \(\PageIndex{26}\) shows the primary cell wall of plants. The primary cell wall contains cellulose microfibrils (no surprise) and two other polymers, pectin and hemicellulose. The middle lamella, consisting of pectins, is somewhat analogous to the extracellular matrix discussed above.
After cell growth, the cell often synthesizes a secondary cell wall, thicker than the primary wall, for greater rigidity. The enzymatic machinery for its synthesis is in the cytoplasm and the cell membrane. It is deposited between the cell membrane and the primary cell wall, as shown in the animated image in Figure \(\PageIndex{27}\).
Figure \(\PageIndex{28}\) shows a structural representation of both the primary and secondary cell walls.
The middle lamella, which contains pectins, lignins, and some proteins, helps "glue together" the primary cell walls of surrounding plants.
Primary Cell Wall:
The main component of the primary plant wall is the homopolymer cellulose (40% -60% mass) in which the glucose monomers are linked β(1→4)-linked into strands that collect into microfibrils through hydrogen bond interactions. Two other groups of polymers, hemicellulose and pectin, make up the plant cell wall.
Hemicellulose can make up to 20-40% of the mass. These polymers have β(1,4) backbones of glucose, mannose, or xylose (called xyloglucans, xylans, mannans, galactomannans, glucomannans, and galactoglucomannans along with some β(1,3 and 1,4)-glucans. The most abundant hemicellulose in higher plants is the xyloglucans with a cellulose backbone linked at O6 to α-D-xylose. Pectin consists of linked galacturonic acids forming homogalacturonans, rhamnogalacturonans, and rhamnogalacturonans II (RGII) [12] [13]. Homogalacturonans (α1→4) linked D-GalA, making up more than 50% of the pectin
Figure \(\PageIndex{29}\) shows some variants of the cell wall components of a plant.
Secondary Cell Wall
The structure of the secondary cell wall depends on the cell's function and environment. It contains cellulose fibers, hemicellulose, and, in addition, a new polymer, lignin. It is abundant in xylem vessels and fiber cells of woody plants. It gives the plant extra stability and new functions, including the transport of fluids through channels.
Lignins, which can make up to 25% of the biomass weight, are made from phenylalanine derivatives but more directly from cinnamic acid. This derivative is made from phenylalanine, which is hydroxylated and converted through other steps to hydroxycinnamyl alcohols called monolignols, as shown in Figure \(\PageIndex{30}\). Three common monomeric (M) derivatives, p-coumaryl, coniferyl, and sinapyl alcohols, can polymerize into lignins, with the units in the polymer (P) named hydroxyphenyl, guaiacyl, and syringyl, respectively.
Lignols are activated phenolic compounds that form phenoxide free radicals (catalyzed by peroxidase enzymes), which can attack other lignols to form covalent dimers. Reaction mechanisms for dimerizing the MS sinapyl alcohol free radical are shown as an example in Figure \(\PageIndex{31}\).
Figure \(\PageIndex{31}\): Dimers of lignols
Now imagine this polymerization continuing, forming additional phenolic free radicals and coupling at many sites to form a large covalent lignin polymer. Figure \(\PageIndex{32}\) shows one example of a larger lignin.
Finally, Figure \(\PageIndex{33}\) shows an image of a poplar tree cell wall, made using surface Raman scattering, showing lignin, cellulose, and lipids in secondary xylem cell walls.
The Extracellular Matrix (ECM) and Basement Membranes
We won't formally discuss cell membranes until Chapter 11, but since anyone reading this book has previously seen biological membranes (including the Gram-negative and positive bilayers discussed above), let's explore a term that most chemistry students, but perhaps not biology students, will find very confusing. That topic is the basement membrane. The basement membrane is encountered so often that we will explore its overall structure here, even though it is not a lipid bilayer. It fits well here, as it is a complex structure composed of proteins and proteoglycans. It's very amorphous, making its structure difficult for those hoping for crystal structures or complex bilayers. It is somewhat similar to the cell wall in functionality. We will offer a cursory explanation. Please visit Introduction to Extracellular Matrix and Cell Adhesion in BioLibre texts for a great overall introduction. Some of the images (when noted) below come from that Cell Biology book chapter.
The extracellular matrix (ECM) is a general term for the large protein and polysaccharide network secreted by some cells in a multicellular organism. They act as connective material to hold cells in a defined space. Cell density can vary greatly between different tissues of an animal, from tightly-packed muscle cells with many direct cell-to-cell contacts to liver tissue, in which some of the cells are only loosely organized, suspended in a web of extracellular matrix, shown in Figure \(\PageIndex{34}\).
The ECM is a generic term encompassing a mixture of polysaccharides and proteins, including collagens, fibronectins, laminins, and proteoglycans, all secreted by the cell. The proportions of these components can vary greatly depending on tissue type. Two quite different examples of ECM are the basement membrane underlying the epidermis of the skin, a thin, almost two-dimensional layer that helps organize skin cells into a nearly impenetrable barrier to many simple biological insults, and the massive three-dimensional matrix surrounding each chondrocyte in cartilaginous tissue. The ability of the cartilage in your knee to withstand the repeated shock of your footsteps is due to the ECM proteins in which the cells are embedded, not to the actual cells that are, instead, few and sparsely distributed. Although both types of ECM share some components in common, they are distinguishable not just in function or appearance but in the proportions and identity of the constituent molecules
Figure \(\PageIndex{35}\) shows a general structure of the basement membrane. Think of it as an amorphous polymer mixture (somewhat similar to a polyacrylamide gel).
Summary
(Summary written by Claude, Sonnet 4.6, Anthropic)
This chapter extends the study of carbohydrates from isolated polysaccharides to their covalent attachments to proteins (glycoproteins and proteoglycans) and their roles as structural polymers in bacterial and plant cell walls and in the extracellular matrix — illustrating the extraordinary information content of glycan structures and the central role of carbohydrates in molecular recognition, structural support, and host-pathogen interactions.
N-linked glycoproteins carry carbohydrate chains attached through GlcNAc to Asn residues within the sequence Asn-X-Ser/Thr (where X is any amino acid except Pro). All three classes — high mannose, complex, and hybrid — share a common pentasaccharide core, (Man)₃(GlcNAc)₂. High-mannose N-glycans extend this core with additional α-linked mannose residues. Complex N-glycans replace terminal mannoses with GlcNAc "antennae" (forming bi-, tri-, or tetraantennary structures), often extended with galactose and capped with negatively charged sialic acid residues. Hybrid N-glycans contain both unsubstituted terminal mannoses and GlcNAc antennae on different branches. The dense glycan coating of viral surface proteins illustrates the functional importance of glycosylation: approximately 50% of the surface area of the SARS-CoV-2 spike protein is covered with biantennary LacNAc N-glycans, which protect underlying protein epitopes from antibody recognition and limit proteolytic degradation. The terminal Gal(α1→3)Gal(β1→4)GlcNAc epitope, present in most mammals (who have an α1,3-galactosyltransferase) but absent in humans, apes, and Old World monkeys (who lost this enzyme through evolutionary mutation), represents the major immunological barrier to xenotransplantation of pig organs into humans; genetic deletion of the pig α1,3-galactosyltransferase gene is a current biotech strategy to make pig hearts immunologically compatible for human transplantation.
The influenza virus provides the most medically consequential example of glycan-mediated molecular recognition. Hemagglutinin (HA), a trimeric glycoprotein on the viral surface, binds sialic acid residues on host cell glycoproteins through its globular head domain (HA1), while the stalk domain (HA2) mediates membrane fusion and viral internalization. Avian influenza viruses prefer Sia(α2,3)Gal (predominant in the avian GI tract), while human pandemic strains H1, H2, and H3 prefer Sia(α2,6)Gal (predominant in the human upper respiratory tract). Swine viruses bind both linkage types, explaining why pigs serve as potential "mixing vessels" for generating human-adapted variants from avian strains. The H5N1 avian flu (which has killed ~56% of the ~240 infected humans as of 2023) retains Sia(α2,3)Gal specificity and targets cells deep in the lower respiratory tract, limiting aerosol transmission. Neuraminidase (NA) on the viral surface is equally essential for the viral lifecycle: it cleaves sialic acid from host cell glycoproteins to release newly budded virions that would otherwise remain tethered to the cell surface. The neuraminidase inhibitors oseltamivir (Tamiflu) and zanamivir (Relenza) block this release step and constitute our primary antiviral defense against influenza.
O-linked glycoproteins carry carbohydrates through N-acetylgalactosamine (GalNAc) attached to Ser or Thr hydroxyl groups without a specific sequon requirement; the position of O-glycosylation is determined by the accessibility of Ser/Thr residues in the folded protein and by the tissue-specific expression of O-GalNAc transferases. Blood group antigens illustrate how small glycan differences determine immunological identity: the ABO blood group system depends on the presence or absence of specific terminal sugar additions to a common H-antigen core oligosaccharide — individuals with blood group A have an N-acetylgalactosaminyltransferase (adding GalNAc), blood group B individuals have a galactosyltransferase (adding Gal), blood group O individuals lack both enzymes and present only the H antigen, and blood group AB individuals have both enzymes.
Proteoglycans represent the extreme of glycoprotein modification, in which proteins can be decorated predominantly with carbohydrate: the protein core carries one or more O-linked glycosaminoglycan chains (heparan sulfate, chondroitin sulfate, keratan sulfate, dermatan sulfate) attached through Ser residues in Ser-Gly repeating sequences, plus variable N-linked oligosaccharides. The variable sulfation patterns along these chains — beyond the control of any genetic template — generate extraordinary structural and informational diversity. Transmembrane proteoglycans like syndecans physically bridge the extracellular matrix and intracellular cytoskeleton, with GAG chains binding matrix proteins (fibronectin, collagen) through basic residue motifs (BBXB or BBBXXB) on the extracellular side and cytoskeletal proteins on the intracellular side. The specific interaction of antithrombin III with a unique pentasaccharide sequence within heparin (rather than with heparin as a whole) exemplifies how sequence-specific GAG recognition mediates precise biological functions.
Bacterial peptidoglycan cell walls are the defining structural element of bacteria and the target of many antibiotics. Both Gram-positive and Gram-negative bacteria build their walls from GlcNAc(β1→4)MurNAc repeating disaccharides (structurally related to chitin, but with a lactate ether group on C3 of every MurNAc); these chains are cross-linked by pentapeptides containing unusual D-amino acids (stabilizing the wall against host proteases that cleave only L-peptides). In Gram-positive bacteria, cross-linking occurs through a pentaglycine bridge between a Lys on one pentapeptide and the terminal D-Ala of another; in Gram-negative bacteria, cross-linking occurs through direct amide bonds involving meso-diaminopimelic acid (m-A₂pm), a meso-compound replacing Lys-3. Remarkably, the entire peptidoglycan sacculus of a bacterium is a single covalent macromolecule — nature's largest single molecule. Gram-negative bacteria additionally possess an outer membrane whose outer leaflet is composed of lipopolysaccharide (LPS); this LPS is synthesized at the inner membrane and transported to the outer membrane by the multi-protein Lpt complex (LptABCDEFG), which uses ATP hydrolysis to extract LPS from the inner membrane and thread it through a periplasmic bridge to the outer membrane. The LPS O-antigen variability determines bacterial antigenicity and serotype, and the outer membrane as a whole is the primary barrier underlying Gram-negative antibiotic resistance. Archaea cell walls differ fundamentally — they contain pseudomurein (using N-acetylalosaminuronic acid instead of MurNAc) or other polysaccharides, and their lipids use ether-linked isoprenoid chains on L-glycerol.
Plant cell walls are among the most complex biological structures known, with ~35 cell types each potentially having distinct wall compositions. Primary cell walls contain cellulose microfibrils (β1→4-glucose, 40–60% of mass, organized through intrachain and interchain hydrogen bonds and hydrophobic stacking into microfibrils), hemicellulose (β1→4-linked backbone polymers of glucose, mannose, or xylose, 20–40%), and pectin (galacturonic acid polymers forming the middle lamella). Secondary cell walls additionally contain lignin (up to 25% of biomass), an amorphous cross-linked polymer formed by oxidative radical coupling of monolignols (p-coumaryl, coniferyl, and sinapyl alcohols derived from phenylalanine) catalyzed by peroxidases; the irregular coupling creates a heterogeneous, chemically resistant polymer that gives wood its strength and recalcitrance to enzymatic degradation. The extracellular matrix (ECM) of animal tissues — composed of collagens, fibronectins, laminins, and proteoglycans in tissue-specific proportions — performs functions analogous to bacterial and plant cell walls, providing structural support and signaling interfaces; the basement membrane, a thin but mechanically important ECM layer underlying epithelial tissues, organizes cells into functional barriers and is the site of cell-ECM communication through integrin receptors.




.png?revision=1&size=bestfit&width=251&height=265)


_from_golden_algae_(9UFE).png?revision=1&size=bestfit&width=402&height=370)
.png?revision=1&size=bestfit&width=199&height=294)
Snip.png?revision=1&size=bestfit&width=147&height=290)