← Back to list

If I were focusing on protein modeling back then…

I might be stumbled into different world of AI nowadays.

Immanuel Sanka · 2025-12-08 05:43 · 92 claps · 8.2 min read
#ai #alphafold #bioinformatics #machine-learning #deepmind
Open on Medium ↗
Wiki topics: ML · Machine Learning AI · AI · General BIN · Bioinformatics PRO · Proteomics & Structure EDU · Education & Learning ⏱️ · Productivity

If I were focusing on protein modeling back then…

I might be stumbled into different world of AI nowadays.

Image was taken from IMDB or here

Image was taken from IMDB or here

These thoughts came to me while watching The Thinking Game from Google DeepMind (the YouTube link can be found **here**) on a casual Thursday night. Before sharing my perspective, I want to revisit how I first entered this field, because it shapes how I see things when watching the documentary.

My life began intersecting with computational biology, especially structural biology, during my master’s studies in 2016. That was when I discovered 3D simulation of DNA, RNA and proteins, and more importantly, their interactions. Before that, most of my work involved small molecules such as carbohydrates, where I focused on predicting inter and intramolecular interactions (more about it **here**). These included hydrogen bonds, covalent bonds and a variety of weaker forces like van der Waals interactions. At that time, my approach was primarily biological, because my chemistry, physics and programming knowledge were almost nonexistent, to be honest. I understood the molecular concepts, but the computational tools were still far outside my comfort zone. Still, that didn’t stop me from doing my thesis project in this field, which involved predicting an unknown 3D protein structure (even without any experimentally determined structure available 😂) directly from a DNA or RNA sequence.

A crystal structure refers to a protein structure solved experimentally at atomic resolution. Historically, X ray crystallography was the main method to obtain such structures. Today, methods such as Cryo EM and NMR spectroscopy are also widely used and, depending on the protein, sometimes preferable. Regardless of the technique, the principle remains the same, we can obtain a highly resolved 3D model from real experimental measurements.

Workflow of 3D structure protein from glutamate receptor-like channel (GLR) (Gangwar et al. 2021)

Workflow of 3D structure protein from glutamate receptor-like channel (GLR) (Gangwar et al. 2021)

Like in the figure above, you can see that the process itself is quite lengthy. It’s just a brief workflow and quick estimate to explain it:

  1. You must express and purify the protein of interest.
  2. Once you have the sample, the next step depends on the equipment available. If you have Cryo EM (like the workflow above), you can begin analysis immediately. If not, you need to go through the more complex process of crystallization.
  3. Crystallization often requires screening hundreds of different conditions (this was exactly the part we struggled with) because even slight changes in pH, temperature or buffer composition can determine whether usable crystals form.
  4. Once crystals are obtained, they are analyzed using techniques such as X ray crystallography or synchrotron X ray diffraction. These experiments generate diffraction patterns or density maps, from which atomic coordinates are derived.
  5. Only after several rounds of refinement does the final 3D model emerge, which is then deposited in the Protein Data Bank.

Using Perplexity, I also tried to sum up the total timeline from start to finish.

Using Perplexity, I also tried to sum up the total timeline from start to finish.

When you look at the whole workflow above, it becomes obvious why each experimentally solved structure represents weeks to months of work and, in the past, even years. This is why it consumes so much time, money and research effort. Because of that, computational modeling and structure prediction have always been attractive for quick hypothesis testing. Even early, imperfect models allowed researchers to explore binding interactions, protein dynamics or structural feasibility before committing to expensive lab work.

Long story short, my interest in protein modeling grew once I realized that prediction tools existed and could generate a hypothetical 3D structures from sequence alone. Using these tools, my supervisors (yes, two supervisors 😄) and I tried to explore what could be achieved computationally, even with limited resources at the time. We tested different software, alignment strategies, compared predicted structures with crystallographic models, examined topology and attempted to validate interactions through docking and spatial analysis. That period opened my eyes to how powerful computational biology could be, especially when combined with good biological intuition. Even though we couldn’t reach strong hypotheses or conclusions, at least I understood that the process would never be simple or quick.

“Good things take time.”

That quote hit me again when I watched The Thinking Game. I almost felt like I jumped into a time machine and wished these tools had existed in 2016. My life would have been easier, probably. Or maybe not, and I will explain why below.

In the documentary, the story became more intriguing when it shifted from the world of games to the world of proteins. What struck me was how naturally that transition happened for Demis Hassabis and the Google DeepMind team. On the surface, games like chess and Go seem completely unrelated to biology. Yet the underlying challenge is the same, a search problem of unimaginable scale where the number of possible states grows far beyond human capability, including protein folding phenomenon. In his Nobel Prize week talk, Demis mentioned three things that makes protein folding problem suits for AI problem solving case, including massive combinatorial search space, a clear objective function and lots of data with an accurate and efficient simulator.

Laureate Demis Hassabis in Nobel talk (source from Youtube or link here)

Laureate Demis Hassabis in Nobel talk (source from Youtube or link here)

Demis and the DeepMind team then tested whether they could help us learn about nature by solving protein folding and simulating 3D structures. Briefly, their first version didn’t meet expectations and the model didn’t really perform well. But once they brought in the right domain expertise (e.g., molecular biologist, computational biologists and bioinformaticians), they improved the model, especially for predicting unknown proteins with no existing homologs. In the documentary, they eventually achieved remarkably high prediction results compared to other computational methods. After finishing the documentary, I started to check and read some publications related to it, especially the CASP 14 related competition report.

From Kryshtafovych et al. (2023), we can also read that “By far the most accurate models are from the AlphaFold group, consistent with the later CASP results. But only one EMA method selected an AlphaFold model as the best, for only one of the targets, and generally the AlphaFold models were not highly ranked.” This raises questions about whether the model still needs refinement. Not because of specificity, but because the known protein structures are probably limited. As of Dec 2025, RCSB hosts 245,663 experimental structures in PDB format and now, we also have 999,251 from AlphaFoldDB. Even though this seems like a lot (you need to see the incremental trend in protein structure release below), they represent diverse and unevenly sampled protein families.

Growth of released protein structure from PDB (per Dec 2025) from here

Growth of released protein structure from PDB (per Dec 2025) from here

During my thesis, I worked with unknown transmembrane proteins, meaning no structural models existed from similar proteins. These proteins are quite difficult to study because they are hydrophobic, unstable outside the membrane and challenging to isolate or crystallize. This is why they are underrepresented in the Protein Data Bank, even though many are key drug targets and essential for cell signaling, transport and communication. When AlphaFold2 (2nd version of AlphaFold) arrived, many researchers were unsure whether it would perform well on transmembrane proteins. The training data contained fewer high quality membrane structures, so expectations were modest. Yet several studies show that AlphaFold2 predicts many transmembrane proteins with accuracy comparable to soluble ones. Some research also reported taht this includes complex multipass proteins, such as ABC transporters, with models often reaching high resolution and low structural errors, RMSD values in the 1 to 2 Ångstrom range and TM scores above 0.9 in blind tests (Hegedus et al., 2022, Tordai et al., 2022, Jambrich et al., 2023). In some comparisons, AlphaFold2 models remained stable during molecular dynamics simulations while older homology models quickly lost their fold. That suggested the predicted structures were not only visually convincing but physically reasonable.

A larger analysis of the human transmembrane proteome showed a similar pattern. AlphaFold2 produces good structures for more than half of all human membrane proteins, especially those with at least some evolutionary or structural information available (Jambrich et al. 2023). The remaining fraction, around 30%, is still difficult to predict. These tend to be very flexible proteins, bitopic receptors or large multichain complexes. AlphaFold2 also predicts a single stable conformation, usually close to an inactive or ground state, so highly dynamic systems like GPCRs or ion channels still require careful interpretation. From EMBL course, we could also see some limitations and things AlphaFold2 struggles to predict.

The limitations of AlphaFold2 from EMBL page

The limitations of AlphaFold2 from EMBL page

Thinking back to my thesis, I think it would still be difficult because no homologs or related structures exist. In my opinion, understudied proteins face the same challenge. I imagine how different the experience would have been if a tool like AlphaFold had existed in 2016. Instead of struggling to build any structural model at all, I would have started with a fairly reliable 3D fold from AlphaFold or similar platform and tried to focus on on understanding its limitations, dynamics and biological implications. However, AlphaFold approaches (versions 1, 2 or 3) did not remove the complexity of protein folding (or interactions for version 3), but they provide an an alternative perspective computationally and give more ideas to be explored in protein folding. I would recommend to read the publications related to AlphaFolds if you need further readings (added in the References section below).

THe newest AlphaFold’s Infrastructure (AlphaFold3 or AF3), taken from Abramson et al. 2024

THe newest AlphaFold’s Infrastructure (AlphaFold3 or AF3), taken from Abramson et al. 2024

I believe the protein folding chapter is still leave mysteries to be explored. The focus could be exploring unknown protein structures, developing better crystallization methods or improving resolution would definitely enhance computational approach. In the future, it possibly help solve these folding mechanisms, one day.

That being said, I feel that this field has been progressive and I am genuinely impressed. Answering the title above, if I had focused on this field, I might have enjoyed playing with tools like AlphaFold. There are also other tools available, such as OpenFold, RoseTTAFold, ESMFold (Meta AI), and many more. I am also cannot wait to see the progress in upcoming years.

I think that’s it for now! 😁

Until my next post!

Reach me on LinkedIn — Immanuel Sanka if you’re interested in a chat! 😄 Let me know! About me? Check my first post here!

References:

  1. Hegedűs T, Geisler M, Lukács GL, and Farkas B. (2022). Ins and outs of AlphaFold2 transmembrane protein structure predictions. Cellular and Molecular Life Sciences, 79(3), 149. https://doi.org/10.1007/s00018-021-04112-1
  2. Tordai H, Suhajda E, Sillitoe I, Nair S, Varadi M and Hegedus T.(2022). Comprehensive collection and prediction of ABC transmembrane protein structures in the AI era of structural biology. International Journal of Molecular Sciences, 23(16), 8877. https://doi.org/10.3390/ijms23168877
  3. Jambrich MA, Tusnady GE, and Dobson L. (2023). How AlphaFold2 shaped the structural coverage of the human transmembrane proteome. Scientific Reports, 13, 18119. https://doi.org/10.1038/s41598-023-47204-7
  4. Kryshtafovych A, Schwede T, Topf M, Fidelis K, Moult J. Critical Assessment of Methods of Protein Structure Prediction (CASP) — Round XIV. Proteins. 2021 Oct 7;89(12):1607–1617. doi:10.1002/prot.26237. PMCID: PMC8726744.
  5. Gangwar SP, Green MN, Yelshanskaya MV, Sobolevsky AI. Purification and cryo-EM structure determination of Arabidopsis thaliana GLR3.4. STAR Protocols. 2021;2(4):100855. doi:10.1016/j.xpro.2021.100855
  6. Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, Ronneberger O, Willmore L, Ballard AJ, Bambrick J, Bodenstein SW, Evans DA, Hung C-C, O’Neill M, Reiman D, Tunyasuvunakool K, Wu Z, Žemgulytė A, Arvaniti E, Beattie C, Bertolli O, Bridgland A, Cherepanov A, Congreve M, Cowen-Rivers AI, Cowie A, Figurnov M, Fuchs FB, Gladman H, Jain R, Khan YA, Low CMR, Perlin K, Potapenko A, Savy P, Singh S, Stecula A, Thillaisundaram A, Tong C, Yakneen S, Zhong ED, Zielinski M, Žídek A, Bapst V, Kohli P, Jaderberg M, Hassabis D, & Jumper JM. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 630(8016), 493–500. https://doi.org/10.1038/s41586-024-07487-w
  7. Nobel Talk from Laureate Demis Hassabis (https://www.youtube.com/watch?v=YtPaZsasmNA)
  8. Protein Data Bank (https://www.rcsb.org/stats/growth/growth-released-structures)

메타데이터
post_id
da0378f2edf9
slug
if-i-were-focusing-on-protein-modeling-back-then-da0378f2edf9
url
https://medium.com/@im-sanka/if-i-were-focusing-on-protein-modeling-back-then-da0378f2edf9
canonical_url
https://medium.com/@im-sanka/if-i-were-focusing-on-protein-modeling-back-then-da0378f2edf9
author_url
https://medium.com/@im-sanka
status
ok
fetched_at
2026-06-21 15:33:18