4.14: Predicting Structure from Sequence and Sequence from Structure/Function (New 10/24)
- Page ID
- 146142
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)(Learning goals written by Claude, Sonnet 4.6, Anthropic)
Predicting 3D Structure from Sequence
- Explain how AlphaFold and RoseTTAFold predict protein 3D structure from primary sequence by integrating multiple sequence alignment (MSA) co-evolutionary information, pairwise residue distance representations, and iterative recycling between sequence and structural information — and interpret structural accuracy metrics (RMSD, TM-score, ipTM, pTM) to critically evaluate the quality of a predicted structure against an experimentally determined one.
- Describe how AlphaFold3 extends structure prediction from single proteins to multi-component complexes, incorporating other proteins, nucleic acids, small-molecule ligands, metal ions, post-translational modifications, and glycans, and use statistical performance data (average RMSD ~1.6 Å, 72% with ipTM > 0.8) to assess both the power and current limitations of complex structure prediction, including outlier cases.
Reverse Protein Folding: From Structure to Sequence and Function
- Explain how FoldSeek's 3Di structural alphabet — encoding the 3D interaction geometry of each residue with its nearest spatial neighbor using 10 features — enables rapid structure-based database searches that circumvent the limitations of sequence-only comparison, particularly for divergent or orphan proteins, and describe how this approach has been used to identify structural and functional homologs among newly predicted viral proteins.
- Describe the diffusion model framework underlying RFDiffusion — in which a forward noising process disrupts a protein structure into random noise and a reverse denoising process regenerates realistic structures guided by deep learning from known structural databases — and explain how conditional inputs (symmetry specifications, binding hotspots, functional motifs, catalytic triads) are used to design proteins with defined structural and functional properties.
- Explain the complementary roles of RFDiffusion (backbone structure generation) and ProteinMPNN (sequence design for a given backbone) in the de novo protein design pipeline, and contrast the data-driven probabilistic approach of ProteinMPNN with physically-based energy minimization methods such as Rosetta in terms of computational tractability and the types of bias each approach introduces.
Applications, Validation, and Remaining Challenges
- Describe specific validated applications of de novo protein design — including IL-6 receptor minibinders for cytokine storm inhibition, universal influenza hemagglutinin binders targeting conserved HA epitopes, symmetric oligomers with defined Cn symmetry, novel fluorescent proteins, and enzymes with computationally positioned catalytic triads — and explain why experimental validation by X-ray crystallography, cryo-EM, NMR, CD spectroscopy, or functional assay remains essential even when predicted and designed structures agree closely.
- Identify the major remaining challenges in computational protein design — designing binders that specifically activate or inhibit protein function, creating efficient catalytic proteins, incorporating conformational flexibility and allostery into designed structures, and engineering novel macromolecular assemblies — and explain why each challenge exceeds what current structure prediction and design tools can reliably accomplish.
Recent Updates: November 2024
The Protein Folding Problem: Sequence to 3D Structure
In Chapter 3.4: Analyses of Protein Structure, we discussed how protein structure can be experimentally analyzed and the 3D structure of a protein can be determined using NMR, X-ray crystallography, and cryo-EM. Now, using sequence and structural databases (such as the PDB, which contains over 227,000 structures), we can often predict a protein's 3D structure from its linear sequence. We can do this by comparing its sequence to homologous proteins (by sequence) whose 3D structures are known. Machine learning and artificial intelligence have extended earlier and simpler "homology modeling" attempts to allow structural predictions for millions of protein sequences using programs such as RoseTTAFold and AlphaFold. RoseTTAFold and AlphaFold produce high-quality structure predictions when trained using the vast sequence information in the Protein Data Bank. Embedded in those linear sequences is a large amount of hidden (to the human eye) evolutionary information that machine learning and AI can harness to predict 3D structures. They work less well when very limited sequence comparisons are available.
In general (for smaller proteins), the protein folding problem, the prediction of structure from sequence, appears to have been "solved". The Nobel Prize in Chemistry in 2024 was awarded to Demis Hassabis and John M. Jumper from Google DeepMind for developing AlphaFold and David Baker for developing RoseTTaFold and other powerful techniques described below. (Check out the previous section, Chapter 4.13: Predicting Structure and Function of Biomolecules Through Natural Language Processing Tools, if you are interested in a "deep dive" into how these programs work!)
A comparison of protein structures obtained with these programs with known 3D structures from X-ray crystallography or other techniques shows they are almost identical. Different metrics can be used to compare predicted structures to the actual ones. The root-mean-square deviation (RMSD) is a common one. RoseTTAFold uses a TM-score to assess the topological similarity of protein structures. Compared to RMDS, the TM-score weights smaller distance errors more than larger ones, making it sensitive to the global fold rather than local structural differences. TM values range from 0 to 100 (100 indicates a perfect match). Scores below 17 indicate no topology match, while those above 50 suggest a common fold.
AlphaFold uses a "neural network, meaning it simultaneously considers patterns in protein sequences, how a protein’s amino acids interact, and its possible three-dimensional structure. In this architecture, one-, two-, and three-dimensional information flows back and forth, allowing the network to collectively reason about the relationship between a protein’s chemical parts and its folded structure". Programs of this type might enable the generation of proteins with new therapeutic or commercial potential from sequences. These include vaccines, sensors, specific immune system suppressors or activators, and antivirals. AlphaFold has now been used to predict the structure of 214 million proteins from more than one million species — essentially all known protein-coding sequences. We have included many AlphaFold iCn3D models throughout this book.
- AlphaFold Database - Protein Structure Database
Figure \(\PageIndex{1}\) shows the backbone tube cartoon of the X-ray PDB structure of the small protein (1xww, cyan) and the structure predicted by both the RoseTTAFold program and AlphaFold (magenta) just from its primary sequence. Sulfate, a competitive inhibitor, is shown (spacefill) bound in the active site. The alignment is spectacular, except for the N-terminal 5 amino acids at the bottom of the figure (6 o'clock). This stretch shows greater disorder, as evident in the X-ray structure, with amino acids exhibiting higher B-factors, indicating greater conformational flexibility.
![]() |
![]() |
Left panel: X-ray structure (cyan) of low molecular weight protein tyrosine phosphatase with bound SO42- (1xww) and corresponding structure predicted by the RoseTTAFold (magenta). Right panel: Same structures using AlphaFold for the structural prediction.
Prediction of Protein-Protein Interactions
AlphaFold3 has also been used to predict the structures of protein complexes, in which multiple copies of the same or different proteins combine to form a larger, quaternary structure. Figure \(\PageIndex{2}\) shows an interactive iCn3D model of a recent stunning example of a predicted AlphaFold complex required to bind a human sperm and egg. The complex consists of three transmembrane human sperm proteins and a human egg protein attached to the egg membrane with a posttranslational lipid anchor (not shown).
Figure \(\PageIndex{2}\): Human sperm proteins and egg protein complex predicted by AlphaFold. (Copyright; author via source). The sequences of the three proteins shown in space fill attached the proteins to the sperm membrane. The egg protein is JUNO (cyan), and the three sperm proteins are 1ZUMO1 (magenta), SPACA6 (brown/gold), and TMEM81 (blue)
Reference for PDB file: Deneke, V. et al. A conserved fertilization complex bridges sperm and egg in vertebrates. Cell. October, 2024. https://doi.org/10.1016/j.cell.2024.09.035. Creative Commons Attribution (CC BY 4.0)
You can download this iCn3D file and load it in iCn3D using these commands to see the structure as rendered in the image above: IMPORTANT: If the file opens as an image in a new browser window, right-click the image and save the file to download it!
- Open iCn3D
- File, Open, iCn3D PNG appendable, and browse for the file in your download folder.
Here is the link to the AlphaFold Server 3. AlphaFold3 is based on a machine-learning process called diffusion, which is explained in more detail below.
The following biological species can be modeled in AlphaFold3:
- macromolecules, including proteins, DNA and RNA
- common ligands including ATP, ADP, AMP, GTP, GDP, FAD, NADP, NADPH, NDP, heme, heme C, myristic acid, oleic acid, palmitic acid, citric acid, chlorophylls A and B, bacteriochlorophylls A and B
- common ions such as Ca2+, Co2+, Cu2+, Fe3+, K+, Mg2+, Mn2+, Na+, Zn2+, and Cl-
- common post-translational modifications (PTMs) of amino acid residues such as phosphorylation of serine, threonine, tyrosine, and histidine, acetylation of lysine residues, methylation of lysine and arginine, malonylation of cysteine, hydroxylation of proline, lysine, and asparagine, palmitoylation of cysteine, succinylation of asparagine, S-nitrosylation, formylation of tryptophan, crotonylation of lysine, citrullination of lysine and arginine
- glycan chains (including branched chains) composed of some sugars, including alpha/beta-D-glucose, alpha/beta-D-mannose, alpha-L-fucose, beta-D-galactose, N-acetyl-beta-D-glucosamine
- common chemical modifications of the DNA (including methylation of cytosine, guanine, and adenine, carboxylation of cytosine, oxidation of guanine, formylation of cytosine) and RNA (including isomerization of uridine into pseudouridine, formylation of cytosine, and methylation of cytosine, guanine, adenine, and uracil
- structures composed of multiple proteins, nucleic acids, ligands, ions, and chemically modified derivatives.
Note that simple drugs are not on the list since those applications are proprietary (at present). AlphaFold 3 can now be downloaded for academic (non-commercial) applications and likely includes drugs. A more limited web version is available here.
It is important to statistically compare the PDB experimentally determined and AF3-predicted structures of complexes. One such comparison involves the wild type and their mutated forms, which should involve just subtle conformational changes. For example, consider human angiogenin and placental ribonuclease inhibitors. Figure \(\PageIndex{3}\) shows an interactive iCn3D model of the experimental human angiogenin and placental ribonuclease inhibitor complex (1A4Y).
Figure \(\PageIndex{3}\): Human angiogenin-placental ribonuclease inhibitor complex (1A4Y). (Copyright; author via source). Click the image for a popup or use this link: https://structure.ncbi.nlm.nih.gov/i...DaUAJzhh8H6LU8
There are 27 mutant variants of the complex whose experimental structures are known. They are included in the SKEMPI database, which contains thermodynamic and kinetic data for wild-type and mutant complexes with known PDB structures. One widely used thermodynamic parameter we have seen before is the change in thermodynamic stability (i.e. ΔΔG0= ΔG0mutant - ΔG0wildtype), in this case for the complex, when key residues lining the binding pocket are mutated. The entire database contains over 317 protein-protein complexes and 8338 mutations. How well does AF3 predict the structure of mutant complexes?
Three statistical values are used to compare the experimental and AF-predicted complex structures:
- RMSD (Root-Mean-Square Deviation) measures the average distance between equivalent atoms in the complex subunits. Lower values show great similarity between the experiential and AlphaFold 3 models.
- ipTM (Interface Predicted Template Model) measures changes in the interface of the subunits between the experimental and AF3 structures. Higher values indicate a closer match.
- pTM (Predicted Template Model), as with simple AF prediction, measures the overall accuracy of the predicted structure (based on both backbone and sidechain orientations). Higher pTM values indicated a more accurate prediction.
Figure \(\PageIndex{4}\) shows the structure of the wild-type RNase inhibitor-Angiogenin complex with 27 red dots indicating mutants with experimental structures (Panel A) and a comparison of the wild-type and a mutant structure (Panel B). Statistical results comparing experimental and AF structures for all 317 complexes in the SKEMPI database are shown in Panels C and D.
Figure \(\PageIndex{4}\): Wildtype and AF predicted composite structure of RNase inhibitor-Angiogenin complexes (Panel A and B) and statistical comparison of all structures in the SKEMPI database (panels C and D). JunJie Wee and Guo-Wei Wei. J. Chem. Inf. Model. 2024, 64, 16, 6676–6683. https://doi.org/10.1021/acs.jcim.4c00976. Published August 8, 2024. CC-BY 4.0 .
Panel A: The cartoon representation of ribonuclease inhibitor-angiogenin complex (PDB ID: 1A4Y). The ribonuclease inhibitor is shown in blue, and the angiogenin is shown in green. 27 mutation spots of 1A4Y in the S8338 data set are indicated in red.
Panel B: The structural alignment of 1A4Y with its AF3 predicted complex.
Panels C and D below are based on the complexes studied, not just the RNase Inhib-Angiogenin complex.
Panel C: The boxplot for RMSD, ipTM and pTM distributions of 317 predicted AF3 protein–protein complexes. RMSDs refer to the overall RMSD calculated by structurally aligning an AF3 complex with its original PDB complex.
Panel D: The breakdown of AF3 protein-protein complexes based on their ipTM and pTM scoring criteria.
The average statistical values for the wildtype and mutant structures are 1.61 Å (RMSD), 0.803 (ipTM), and 0.847 (pTM), respectively. The graphs clearly show that most predictions (72%) had high ipTM scores (> 0.8), while 99% had reasonably high pTM scores (>0.5). However, they were outliers, as shown in Panel C, with RMSD values > 4 Å, indicating poorer performance with AF3. Most complexes with high prediction values also had low RMSD values. Experiments like these will be used to continually refine programs such as AlphaFold 3.
Reverse Protein Folding Problem: 3D Structure to Function
So now we have structures of 200 million plus proteins. Pick one, Protein X, of unknown function. What might its function be? Surely, its sequence could be compared to the entire database to find homologous proteins that might give a clue to the function of Protein X. But what if the sequence of Protein X is very divergent from potential sequence homologs, since they are very distant from each other evolutionarily? Also, what if comparison proteins in the database are underrepresented? For example, our knowledge of the sequences and structures of proteins from pathogens (viruses and bacteria) is incomplete, especially since we have studied just a small fraction of the virus and bacterial world.
To circumvent this problem, the actual predicted or determined 3D structure of Protein X (not its sequence) could be compared with 3D structures in databases. This should be difficult, since it would be a 3D comparison rather than a 1D comparison of linear sequences. A program called FoldSeek makes 3D structural comparisons computationally easier.
Instead of using an alphabet of actual sequences (such as the single-letter code for the 20 naturally occurring amino acids - ACDEFGHIKLMNPQRSTWYV), a new "structural alphabet" based on the conformations of short stretches of 3-5 alpha C (Cα) atoms in the protein backbone has been used, but this doesn't explicitly contain tertiary interactions found in proteins. Instead, FoldSeek uses a 3D interaction alphabet (3Di) with 20 states (one for each amino acid), each with 10 interaction "features." A conformational state for residue X is defined for its closest spatial residue, y. The state description is less dependent on the next amino acid in the linear sequence for a given amino acid. The defined state has more information when x is in a conserved and packed protein core than in a nonconserved, more flexible loop. In contrast, there would be less information if just the backbone structural alphabet were used. Figure \(\PageIndex{5}\) below gives a pictorial view of how the 3Di state for a single amino acid Val at a specific position in a 3D structure is defined. Note that the state has 10 3D features, more than just the conformation of a backbone of 3 amino acids or the next amino acid in the linear sequence.
Figure \(\PageIndex{5}\): Learning the 3Di alphabet. van Kempen, M., Kim, S.S., Tumescheit, C. et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol 42, 243–246 (2024). https://doi.org/10.1038/s41587-023-01773-0. Creative Commons Attribution 4.0 International License. http://creativecommons.org/licenses/by/4.0/.
(1) 3Di states describe the tertiary interactions between a residue i and its nearest neighbor j. Nearest neighbors have the closest virtual center distance (yellow). Virtual center positions were optimized for maximum search sensitivity. (2) To describe the interaction geometry of residues i and j, we extract seven angles, the Euclidean Cα distance, and two sequence distance features from the six Cα coordinates of the two backbone fragments (blue and red). (3) These 10 features are used to define 20 3Di states by training a VQ-VAE modified to learn states that are maximally evolutionarily conserved. The encoder predicts the best-matching 3Di state for each residue during structure searches.
Recently, FoldSeek has been used to identify structural and functional similarities between over 67,000 newly predicted viral proteins (underrepresented in the PDB) and other proteins of known structure. Of these:
- 62% had distinct structures and were not homologous to proteins in the AlphaFold database (as we indicated above).
- Many of the 38% left were structurally homologous to nonviral proteins, suggesting that viral proteins share functional similarities with host analogs.
Similar 3D structures imply similar functions, so probable functions could be described for some novel proteins. Some were involved in the viral escape from the host's innate immune system. We'll explore that in Chapter 5.4: Complementary Interactions between Proteins and Ligands - The Immune System.
The FoldSeek server allows multi-database searches, including AlphaFoldDB (version 4: Proteomes and Swiss-Prot), AlphaFoldDB (version 4) and CATH25 clustered at 50% sequence identity, ESM Atlas-HQ and Protein Data Bank (PDB). These Google Colab sites are also available:
In summary, FoldSeek is useful in several circumstances. You ...
- have a protein sequence, but a comparison to other sequences doesn't give you enough information. A structure-based search would then be helpful.
- want to design a brand new protein (de novo protein synthesis), and you would want to know that its structure is not similar to other proteins
- want to design a protein with a particular function, and want to compare its structure to other proteins with a similar function
Reverse Protein Folding Problem: 3D Shape to Linear Sequence - Designing Proteins From Scratch
In yet another use of machine learning and artificial intelligence, programs can start with a desired 3D shape (a protein backbone, for example) and determine the amino acid sequence needed to achieve it. Two programs, ProteinMPNN and RoseTTAFold Diffusion (RFDiffusion), developed by David Baker (who also won the Nobel Prize in Chemistry in 2024) et al, have enabled these predictions. It allows protein structure design, not structure prediction.
Yet another dream that seemed so distant not so long ago was to design a protein from scratch with no linear sequence (hence little alignment information) but with a final desired structure or function in mind. Here are some possible "de novo" design examples of novel proteins that...
- are soluble versions of a known membrane protein, which could advance drug design;
- bind with high affinity to a desired small molecule (much like an antibody), enabling the creation of sensors and protective agents;
- bind target molecules and catalyze their chemical conversion to products, allowing the development of new and nontoxic catalysts;
- bind to another target protein and modulate its function by activating or inhibiting it;
- have novel, unrepresented folds that could further elucidate key principles of protein folding and stability while creating new functionalities.
RFDiffusion
This dream has also been accomplished in large measure. David Baker is a pioneer in de novo protein structure design and prediction. His group has developed and used several programs, including RoseTTAFold Diffusion (RFDifffusion), which uses AI to design new proteins with novel structures and functions. RFDiffusion is freely available to anyone for use in Google Collaboratory. It creates new structures by combining structure prediction from RoseTTAFold with an AI "Diffusion" model.
To understand the term diffusion in structure prediction, let's first explore AI/machine learning models for generative image creation. Instead of starting with no previous information, start with a clear image, add random (Gaussian) noise to it (noising), and then try to recreate the original image by a "denoising" process. Some prior information and additional programs would help constrain the denoising process for generating a requested image. This process is illustrated in Figure \(\PageIndex{6}\) below. Note that the arrows are reversible.
Figure \(\PageIndex{6}\): Xiao, H., Wang, X., Wang, J. et al. Single image super-resolution with denoising diffusion GANS. Sci Rep 14, 4272 (2024). https://doi.org/10.1038/s41598-024-52370-3. Creative Commons Attribution 4.0 International License. http://creativecommons.org/licenses/by/4.0/.
Simplistically, this is similar to solving protein X-ray structures. A given protein in a crystal produces an X-ray diffraction pattern specific to the atoms and their arrangement in the crystal lattice. In the reverse process, the X-ray diffraction pattern can be computationally analyzed to produce the arrangement of atoms (from an initial electron density map) in the lattice that would generate the given diffraction pattern.
In a diffusion model for generating protein structures using RFDiffusion, randomly disordered small chemical fragments diffuse together to form a more ordered, realistic protein structure. Information from the database of known protein structures is used to constrain the generative processes through a deep learning–based protein sequence design method called ProteinMPNN (Protein Message-Passing Neural Network). It differs from Rosetta, a physically based method that maximizes side-chain packing to reach the lowest-energy state. Designing a sequence that produces the lowest energy state is more computationally challenging than finding it for a given sequence. Calculating energies of unwanted nonproductive oligomeric and aggregated states makes this approach intractable.
What is more doable is to carry out these two steps in succession:
- first, search for the lowest-energy sequence for a given backbone structure;
- then search the "universe" of possible structures for the sequence created in the first step to determine if it is indeed the lowest energy conformation.
Methods like Rosetta use physical "rules" to minimize undesired results. For example, restrictions are used to place hydrophobic side chains on the surface of a protein, as these might promote unwanted aggregation. ProteinMPNN addresses these issues by using data from all solved structures to estimate the most probable amino acid at each position. It requires less human theoretical knowledge as it extracts an energy-minimized folded state from an immense amount of structural data. It's a bit like deriving Newton's Laws of Motion from data without the underlying theory, even though the data was acquired from systems whose motions and positions are well described by Newton's Laws.
Figure \(\PageIndex{7}\) below shows the noising (right to left) and denoising (left to right) processes that can generate a protein structure in a diffusion model. It parallels the image deconstruction and reconstruction shown in Figure \(\PageIndex{6}\) above.
Figure \(\PageIndex{7}\): Protein design using RFdiffusion. Diffusion models for proteins are trained to recover corrupted (noised) protein structures and to generate new structures by reversing the corruption process through iterative denoising of initially random noise XT into a realistic structure X0 (top panel). Watson, J.L., Juergens, D., Bennett, N.R. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023). https://doi.org/10.1038/s41586-023-06415-8. Creative Commons Attribution 4.0 International License. http://creativecommons.org/licenses/by/4.0/.
RFDiffusion can generate a protein sequence without imposing constraints on the final structure (an unconditional process). Conditions placed on the denoising process can lead to conditioned structures. Conditions such as symmetric noise, a binding target, a functional motif such as pre-positioned amino acids in an active site, and a symmetric motif can lead to synthetic oligomers, a binder protein that interacts with a target protein, an active site with the correct 3D disposition of catalytic residues, and symmetrical scaffolds, respectively. These examples are illustrated in Figure \(\PageIndex{8}\) below.
Figure \(\PageIndex{8}\): b, RFdiffusion is broadly applicable for protein design. RFdiffusion generates protein structures without further input (top row) or by conditioning on (top to bottom): symmetry specifications; binding targets; protein functional motifs, or symmetric functional motifs. In each case, random noise and conditioning information are input to RFdiffusion, which iteratively refines that noise until a final protein structure is designed. Watson, J.L. et al., ibid.
Figure \(\PageIndex{9}\) below shows that the final structure predicted for a 300 amino acid protein sequence by AlphaFold (bottom row) is almost identical to the final structure produced by RFDiffusion (top row).
Figure \(\PageIndex{9}\): An example of an unconditional design trajectory for a 300-residue chain, depicting the input to the model (Xt) and the corresponding X^0 prediction. At early timesteps (high t), X^0 bears little resemblance to a protein but is gradually refined into a realistic protein structure. Watson, J.L. et al., ibid.
Unconditional RFDiffusion models for protein sequences up to 600 amino acids are essentially identical to AlphaFold's.
Figure \(\PageIndex{10}\) below shows how hot spots (key binding residues) in a target protein are used as a condition in an RFDiffusion model in the de novo synthesis of a mini-binder for a target protein.
Figure \(\PageIndex{10}\): RFdiffusion generates protein binders given a target and specification of interface hotspot residues. Watson, J.L. et al., ibid.
The video below from the Baker lab (obtained from YouTube at https://youtu.be/geqlzPsigQo) shows how a protein structure can be created that binds to a predefined structure, in this case, the insulin receptor, using conditional RFDiffusion.
The same video is found at the Baker site at: https://www.bakerlab.org/2023/07/11/...rotein-design/
The structures produced by RFDiffusion and ProteinMPNN for any given sequence can be verified by making the protein and analyzing its structure using X-ray crystallography, NMR, or cryoEM. Less precise methods, such as CD spectroscopy, are also used to obtain a simpler measure of the overall secondary structure of the synthesized protein.
Examples
The following iCn3D models of crystal structures using these methods illustrate the power of RFDiffusion methods in creating new proteins of defined structure and function.
| Examples of de novo protein design | Interactive iCn3D model with links |
|
GP130 (IL6 coreceptor) in complex with a de novo designed IL-6 mimetic (8UPA) Cytokine storms (cytokine release syndrome) are often deadly inflammatory responses that accompany bacterial or viral infections (such as COVID-19 in those with severe disease). The storm is associated with the overexpression and release of two proinflammatory cytokines, interleukin 1 (IL-1) and interleukin 6 (IL-6), by activated immune cells such as macrophages. IL-1 and IL-6 bind to their receptors (IL-1R and IL-2R) with high affinity. IL-6 also binds to a coreceptor (GP130) needed for cytokine release. Inhibitors used to interfere with the cytokines:receptor complex can persist too long and have deleterious effects, since an appropriately amplified immune response is needed against bacteria and viral infections. A small protein antagonist (a minibinder or MB) with a high affinity (pM to nM dissociation constant) for the receptor and the IL-6 coreceptor was made through de novo design and proved protective against a cytokine storm in animal models. The structure of the IL-6 mimetic with its coreceptor, GP130, is shown in the iCn3D model to the right. The computationally designed structure was a close match to the X-ray structure. Reference: Huang, B., Coventry, B., Borowska, M.T. et al. De novo design of miniprotein antagonists of cytokine storm inducers. Nat Commun 15, 7064 (2024). https://doi.org/10.1038/s41467-024-50919-4 |
Here is another link to see a surface representation of the two interacting proteins: https://structure.ncbi.nlm.nih.gov/i...3bbFWhy2eHNGAA |
|
Designed Influenza HA binder, HA_20, bound to Influenza HA (8SK7) Hemagglutinin (HA) from the influenza virus is a trimeric membrane protein. Each "monomer" is a heterodimer consisting of two different chains, HA1 and HA2. The HA1 domain binds to a specific sugar, sialic acid, found on many human cells, particularly in the respiratory tract. The HA2 subunit is transmembrane. We receive a vaccine each year that targets the HA protein, since the globular head of HA, which interacts with the human cell surface, mutates so quickly from year to year. Large structural shifts in the influenza HA protein can trigger pandemics. Parts of the HA molecule are more conserved and are somewhat sequestered from the human immune system. Targeting them could lead to a more permanent and universal vaccine. A small influenza binder was synthesized de novo and tightly bound (nanomolar dissociation constant). The de novo-designed protein had essentially the same structure as the AlphaFold computational model. An iCn3D model showing the interaction of the HA binding with one HA1:HA2 heterodimer is shown to the right. Reference: Nature 620, 1089–1100 (2023). https://doi.org/10.1038/s41586-023-06415-8. Creative Commons Attribution 4.0. International License. http://creativecommons.org/licenses/by/4.0/. |
The gray is HA2, the cyan is the HA1, and the red/yellow-coded secondary structure is the designed HA minibinder. The biological HA complex contains three copies of the heterodimeric structure shown above Here is another link to see a surface representation of the three interacting proteins: |
|
Pentameric helical bundle protein (8U5W) De Novo protein synthesis was used to create a single protein chain (i.e., a monomeric protein with a single C5 rotational symmetry axis. Open the iCn3D model to the right. It contains a single rotational axis (red line). Rotation around the axis by 3600/5 reproduces the identical structure. The designed protein also displays near-infrared fluorescence upon binding to the synthetic dye merocyanine. The protein forms a covalent Schiff base with the dye. If the Schiff base is protonated, the fluorescence spectra show a large red shift in both the excitation and emission wavelengths. The protein/dye complex can be used for tissue imaging at greater depths than other visible-light fluorophores. Reference: https://www.researchsquare.com/article/rs-4652998/v1 |
|
|
Symmetric Oligomers In contrast to the previous example of a symmetric monomer, the de novo protein models to the right contain multiple subunits in oligomers that display different types of cyclic symmetry. Figure \(\PageIndex{14}\) to the right displays C2 symmetry, with a rotation of 3600/2 around the axis resulting in an identical structure. The dimer also displays allostery: it changes its shape globally upon the addition of effector molecules. Allostery will be explained in Chapter 5. Figure \(\PageIndex{15}\) to the right is a homo 6-mer of identical subunits and C6 symmetry. Rotation of 3600/6 around the axis results in an identical structure. Figure \(\PageIndex{16}\) to the right is a homo 8-mer of identical subunits and C8 symmetry. Rotation of 3600/8 around the axis results in an identical structure.
|
|
|
A protein with an active site A protein was designed using RFDiffusion to recreate an active site containing three key catalytic residues from the native enzyme, cytotoxic ribonuclease alpha-sarcin (1DE3). The left model in Figure \(\PageIndex{17}\) below and the iCn3D model in Figure \(\PageIndex{18}\) in the adjacent right column show the three active site residues used for conditional de novo protein modeling. The middle two images below show the isolated catalytic "triad" input and the structure generated by RFDiffusion. The right image below is a zoomed image of the active site in the designed protein.
Figure \(\PageIndex{17}\): Comparison of native ribonuclease sarcin and RFDiffusion designed protein. Supplemental Figure, Watson, J.L. et al., ibid. |
This iCn3D model is for cytotoxic ribonuclease alpha-sarcin (1DE3). Three active site amino acids, H50, E96, and H137, were conditionally used to create the de novo protein with the same active site residues (shown to the left).
|
In Chapter 11.1, we will explore how RFDiffusion can create novel membrane proteins and soluble versions.
One final comment: Structures predicted by these AI programs must be subjected to experimental validation of structure and function. Since creating new structures with designed functions is so easy, we must be careful not to blindly accept the results without supporting experimental validation.
AlphaProteo
AlphaProteo from Google DeepMind is also used to design protein binders targeting specific protein sites. Download this file for a video of a synthesized protein binder designed for the SARS-CoV-2 spike receptor-binding domain (reference). This program is not yet available for use (as of 11/11/24). The machine learning methods used in AlphaProteo were not reported in the preprint reference because of "biosecurity and commercial considerations," so we can't explain the basis of the program as we did above for RFDiffusion. Figure \(\PageIndex{19}\) below shows, in general, the steps involved in developing binders that interact with "hotspots" sites on target proteins.
Figure \(\PageIndex{19}\): Overview and experimental performance of AlphaProteo. Vinicius Zambaldi et al. De novo design of high-affinity protein binders with AlphaProteo. Submitted 9/12/24. https://arxiv.org/abs/2409.08022. https://creativecommons.org/licenses/by-nc-sa/4.0/
Panel (A) Schematic of the design system. The generative model outputs designed structures and sequences of binder candidates, and the filter is a model or procedure that predicts whether a design will bind.
Panel (B) Schematic of target-structure-conditioned binder design as performed by the generative model.
Panel (C) Crystal structures (light yellow) and hotspot residues (dark yellow spheres) of seven target proteins for binder design experiments in this work. VEGF-A and IL-17A are both disulfide-linked homodimers. See Table S1 for PDB IDs and hotspot residue numbers.
Figure \(\PageIndex{20}\) below shows the interactions of the de novo synthesized binder with seven target proteins.
| no binder reported |
Figure \(\PageIndex{20}\): Biochemical characterization of representative binders for each target-design model.
The binders all interacted tightly with their target protein.
Challenges that remain
Here are some examples that pose challenges
- Binders that affect the function of a protein: These include both small-molecule binders (i.e., drugs) that target orthosteric or allosteric sites. Essentially, this is the task of the drug and pharmaceutical industry. Designing binders is especially hard for membrane proteins. Also, binders that mimic small drugs are difficult to obtain, since databases are more limited and often proprietary, so the training set is smaller. In addition, the differences between binders that activate or inhibit a target protein can be subtle.
- de novo synthesis of protein catalyst: Much of a protein structure is used to bring key groups into a stable configuration for catalysis. Synthetic chemists aim to make small transition-metal catalysts that mimic the function of proteins, with a catalytic site that often contains a metal ion. This suggests that natural proteins might not be the most efficient mimics to produce novel protein catalysts. Also, proteins that differ in 3D structure can carry out similar reactions.
- Conformational flexibility in proteins: Unless we look at the dynamic structures of proteins, our minds can be trapped into creating just the most stable, low-energy protein structure. Yet flexibility and conformational changes are key to protein function and regulation. Programming conformational change into de novo synthesis algorithms is another complex task.
- Creating proteins and protein complexes with functions other than catalysis: Many macromolecular assemblies (inflammasomes, proteasomes, regulated membrane pores, mobility proteins, etc) provide critical cellular functions. Creating new ones could offer novel ways to modulate cell function. One example would be to create nanoparticles that can deliver "cargo" (such as vaccines) into cells or sequester and eliminate deleterious intracellular components (such as misfolded and aggregated proteins).
Summary
(Summary written by Claude, Sonnet 4.6, Anthropic)
This chapter surveys the current state of computational protein structure prediction and de novo design — arguably the most transformative development in structural biochemistry in decades — moving from the well-established "forward problem" (sequence to structure) through the "reverse problem" (structure to sequence and function) and concluding with real-world examples of computationally designed proteins with validated biological activity.
Sequence-to-structure prediction has been solved effectively for most small- to medium-sized globular proteins. AlphaFold and RoseTTAFold achieve near-atomic accuracy by exploiting the evolutionary information embedded in multiple sequence alignments: when two residues that are spatially close in the 3D structure coevolve (because mutations in one residue must be compensated by mutations in the other to preserve fitness), pairwise co-mutation patterns in a large MSA reveal which residue pairs are spatially proximal. AlphaFold's EvoFormer processes this MSA using row-wise and column-wise attention transformers (T1) while simultaneously building a pairwise distance representation of the single input sequence via triangular self-attention (T2). Its structure module integrates both to predict atomic coordinates. A recycling strategy in which structural information informs the MSA representation and vice versa dramatically improves accuracy. Structural quality is assessed by RMSD (average deviation of equivalent atoms), TM-score (global topology similarity), and for complexes, the interface-specific ipTM and global pTM metrics. AlphaFold3 has generalized the framework to predict the structures of multi-component assemblies, including proteins, DNA, RNA, ions, common metabolites, post-translational modifications, and glycan chains, achieving average RMSD values of ~1.6 Å for the 317 protein-protein complexes in the SKEMPI database, with 72% achieving ipTM > 0.8 and 99% achieving pTM > 0.5 — though outliers with RMSD > 4 Å indicate ongoing limitations for some complexes. AlphaFold3 for academic non-commercial use and AlphaFold for 214 million predicted protein structures across more than one million species are now publicly accessible.
Structure-to-function prediction via FoldSeek addresses the reverse problem: given a protein of unknown function, find structurally similar proteins in databases to infer probable function. FoldSeek achieves this without computationally expensive 3D superposition by defining a "3Di" structural alphabet of 20 states, each encoding the 3D interaction geometry between a residue and its nearest spatial neighbor using 10 features (seven inter-residue angles, a Cα distance, and two sequence-distance features). This structural alphabet is more information-rich than backbone conformation alone because it captures tertiary interactions, and it enables sequence-alignment-like database searches at the structural level. Applied to 67,000 newly predicted viral proteins, 38% showed structural homology to known non-viral proteins, revealing probable functions including roles in viral immune evasion. FoldSeek is particularly valuable for divergent proteins where sequence identity is too low for homology to be detected by conventional sequence alignment.
De novo protein design — creating entirely new proteins with desired structure and function — has been accomplished using two complementary tools developed by David Baker's group, contributions for which Baker shared the 2024 Nobel Prize in Chemistry with AlphaFold developers Jumper and Hassabis. RFDiffusion is a generative AI model that creates protein backbone structures by reversing a diffusion process: during training, known protein structures are progressively noised (randomized) and the network learns to denoise them back to realistic structures; at design time, purely random noise is iteratively denoised into a plausible protein backbone, guided by conditioning inputs that can specify symmetry (generating Cn oligomers such as C2, C6, C8, or C5 monomers), protein binding targets and hotspot residues, functional motifs such as pre-positioned catalytic residue triads, or combinations thereof. ProteinMPNN then identifies the optimal amino acid sequence for the RFDiffusion-generated backbone by drawing probabilistically from the patterns in all solved structures rather than applying physical energy functions — a data-driven approach that avoids the computational intractability of calculating energies for all possible misfolded and aggregated alternative states. Structures designed by this pipeline are validated by expressing the designed sequence, then determining the actual structure by X-ray crystallography, NMR, or cryo-EM and comparing it to the design model.
Validated examples span a remarkable range: a de novo IL-6 receptor minibinder that protects against cytokine storm in animal models; an influenza hemagglutinin minibinder targeting conserved HA epitopes to potentially serve as a universal influenza antigen; a pentameric helical bundle with near-infrared fluorescence enabling deep-tissue imaging; symmetric protein oligomers with programmable Cn symmetry and, in one case, allosteric conformational switching; and a novel protein scaffold containing the same three-residue catalytic triad geometry as a natural cytotoxic ribonuclease. In all cases, the designed structures closely matched both the computational model and the AlphaFold prediction of the designed sequence, validating the pipeline and demonstrating its potential to generate proteins with new therapeutic, industrial, and scientific applications.
Significant challenges remain. Designing proteins that specifically activate or inhibit a target protein's function (not merely bind it) requires understanding subtle structural differences between agonist and antagonist binding geometries. Creating efficient de novo catalysts faces the challenge that natural enzymes use their entire structure to precisely position a relatively small active site. Incorporating conformational flexibility and allostery — critical for regulated protein function — into design algorithms requires moving beyond static energy minima toward dynamic ensembles. And constructing functional macromolecular assemblies such as regulated pores, degradation machines, or delivery nanoparticles remains largely beyond current capabilities. Crucially, computational predictions are hypotheses, not facts: experimental validation of both structure and function is always necessary before any computationally designed protein is accepted as achieving its intended purpose.







.png?revision=1&size=bestfit&width=481&height=383)
_in_complex_with_a_de_novo_designed_IL-6_mimetic_(8UPA).png?revision=1&size=bestfit&width=445&height=299)

.png?revision=1&size=bestfit&width=238&height=219)




.png?revision=1&size=bestfit&width=303&height=211)