← Back to list

AI Approaches in PROTACs: The New Frontier in Targeted Protein Degradation

In continuation of our previous blog series on PROTAC Design, Parts 1 and 2 of this series set the stage: Part…

Dr. Riyaz Syed · 2025-10-29 12:16 · 0 claps · 8.5 min read
#protac #ai-approaches-in-protacs #protein-degradation #ai-models-in-protacs
Open on Medium ↗

AI Approaches in PROTACs: The New Frontier in Targeted Protein Degradation

In continuation of our previous blog series on PROTAC Design, Parts 1 and 2 of this series set the stage: Part 1(https://medium.com/@riyaz_76592/unlocking-the-potential-of-protacs-drugging-the-undruggable-with-targeted-protein-degradation-fead042470e2) described the medicinal-chemistry 3L framework (Ligand, Linker, Ligase); Part 2 (https://medium.com/@riyaz_76592/rational-protac-development-integrating-docking-md-and-energetics-for-efficient-drug-discovery-0ba632d021c6) described the physics-based computational stack — docking, ternary modeling, molecular dynamics (MD), and free-energy calculations used to triage designs before synthesis. While these tools are powerful, they are expensive, resource intensive and often require careful handcrafting for each target system.

Over the past five years, artificial intelligence (AI) — from graph neural networks to diffusion and reinforcement-learning models has begun to transform PROTAC design. AI excels at learning subtle, nonlinear structure — activity relationships, proposing chemically tractable linkers, and, predicting degradation outcomes that depend on three-body cooperativity.

The final part of this blog series explains how AI can augment and accelerates the PROTAC pipeline in conjugation with computational and medicinal chemist, what it can (and cannot) do today, and where the field is heading.

Figure 1: AI/ML in PROTAC Design (Gharbi Y. et al., Digital Discovery, 2024, 3)

Figure 1: AI/ML in PROTAC Design (Gharbi Y. et al., Digital Discovery, 2024, 3)

The PROTAC Challenge: Why Traditional Methods Fall Short

PROTACs are composed of three components: a ligand binding the target protein, a linker, and an E3 ligase recruiter. This “3L framework” generates a design space so vast that conventional approaches struggle:

  • Hundreds of known target ligands
  • Thousands of possible linker chemistries
  • Dozens of E3 ligase recruiters

This yields millions of theoretical PROTACs. Physics-based methods like molecular dynamics (MD) provide accurate modeling but screening the entire space would take decades as it consumes an enormous amount of computational time and resources. AI, designed to learn patterns from data, offers efficient exploration and generation of novel solutions within these constraints.

The Rise of Data-Driven PROTAC Discovery

The PROTAC field has expanded rapidly: in 2015, only a few dozen bifunctional degraders were reported. By 2025, public repositories list over 6,000 PROTAC molecules, spanning hundreds of E3 ligases and warheads.

Core public resources driving AI integration:

  • PROTAC-DB 3.0 — now hosts ~6,111 PROTACs, 569 warheads, 2,753 linkers, and 107 E3 ligases. It provides physicochemical and pharmacokinetic parameters, experimental degradation data (DC₅₀, Dmax), and even predicted ternary complex structures for active compounds [1]
  • PROTACpedia — curates ~1,190 experimentally validated PROTACs covering 105 targets and 82 PDB structures. It emphasizes quality-controlled degradation assays and binding modes, serving as a high-fidelity source for ML benchmarking [[2]](http://2. https://protacpedia.weizmann.ac.il/)

These curated datasets form the substrate for machine-learning pipelines that extract latent chemical rules beyond human intuition or classical QSAR.

AI in the PROTAC Workflow: From Design to Degradation

AI models can be embedded at nearly every stage of degrader design, from fragment selection to ternary-complex prediction and degradation outcome forecasting.

1. AI-Driven Ligand and Warhead Discovery

At the entry point of the 3L framework, AI accelerates the identification of both protein-of-interest (POI) ligands and E3 ligase binders.

  • Graph Neural Networks (GNNs) and transformer-based molecular encoders (e.g., ChemBERTa, DeepChem) learn molecular embeddings directly from SMILES or 3D graphs to predict binding affinities and selectivity profiles [3].
  • Transfer learning on biochemical datasets like MoleculeNet allows models to generalize across protein families **[4]**.
  • Integrating AlphaFold-Multimer structural fingerprints with learned embeddings enables efficient virtual screening against novel targets **[5]**.

Though warhead AI is maturing, the biggest differentiator lies in linker and degradation modeling, as detailed below.

2. AI-Based Linker Design and Enumeration

The linker remains the most delicate and important component of PROTAC design, it governs ternary-complex geometry, cooperative binding, and cell permeability. Traditionally, medicinal chemists relied on intuition or brute-force enumeration; AI now offers structured, generative alternatives.

Table 1: Key models and tools:

Figure 2: The PROTACs and generated linkers for BRD4-PROTAC-VHL by DiffPROTACs. (Li F. et al. Briefings in Bioinformatics, 2024, 25(5), bbae35)

Figure 2: The PROTACs and generated linkers for BRD4-PROTAC-VHL by DiffPROTACs. (Li F. et al. Briefings in Bioinformatics, 2024, 25(5), bbae35)

Emerging hybrid workflows:

  • Combining DiffLinker + Docking/MD refinement yields hundreds of plausible linkers filtered by ternary-complex stability and solubility metrics.
  • Active-learning frameworks iteratively retrain on successful designs, continuously improving generative fidelity.

3. Predicting Ternary-Complex Formation and Cooperativity

A unique challenge in PROTACs is modeling three-body interactions — protein of interest (POI), E3 ligase, and degrader. AI provides an efficient complement to exhaustive MD simulations.

  • Graph- and ML-based scoring architectures have been developed to map interface features from PROTAC-induced ternary complexes and to estimate cooperativity, providing a quantitative understanding of favorable bridging geometries [10, 11].
  • PROTAC-RL introduced a reinforcement-learning framework that optimizes linker composition and conformation based on predicted ΔG binding and degradation efficiency [12].
  • AlphaFold-Multimer + ML refinement now generates ternary-complex hypotheses consistent with cryo-EM-resolved degrader complexes, closing the design–validation loop [13, 14].

These systems accelerate design cycles by orders of magnitude while providing interpretable insights into cooperativity drivers.

Table 2 — Model comparison (Linker generation & ternary scoring)

4. AI-Enhanced Degradation Prediction

Once a PROTAC binds, the next question is: will it actually degrade the protein inside the cell? This step involves complex interplay among permeability, ternary half-life, and ubiquitination efficiency.

AI models address this multivariate problem directly.

  • DeepPROTACs, trained on > 3,000 experimentally characterized degraders, predicts degradation potency (DC₅₀ and Dₘₐₓ) using combined molecular and structural embeddings. It achieves ~77.95 % prediction accuracy and AUROC ~0.847 on a test set. The dataset was curated from PROTAC-DB and labeled according to DC₅₀ and Dmax values. This model demonstrated superior performance compared to traditional QSAR models, highlighting its potential in guiding the design of effective PROTACs [15].
  • Accelerated Rational PROTAC Design for BRD4 employed a hybrid AI-physics approach to design 5,000 candidate PROTACs targeting BRD4. These compounds were filtered using machine learning classifiers and molecular simulations. Six were synthesized, and three demonstrated effective degradation of BRD4 in cellular assays, with one showing favorable pharmacokinetics in mice [12].

Figure 3: AI-designed BRD4–CRBN PROTAC generated using the PROTAC-RL model and predicted degradation activity

Figure 3: AI-designed BRD4–CRBN PROTAC generated using the PROTAC-RL model and predicted degradation activity

  • Modeling PROTAC Degradation Activity (Ribes et al.) report that an ensemble of three ML models (combining deep learning and gradient boosting) reaches ~82.6% accuracy (AUC 0.848) on PROTAC degradation prediction, “comparable to state-of-the-art” methods. These results indicate that data-driven ML approaches on PROTAC-DB surpass earlier QSAR-like methods for distinguishing degraders from non-degraders [16].
  • Large-language-model-based property predictors (e.g., MolGPT, ChemLLM) enable prompt-guided molecular optimization to design PROTACs with tailored potency and E3 selectivity [17].

Synergy Between Physics and Machine Learning

AI does not replace physics-based methods it extends them. Hybrid physics-guided ML is now the dominant paradigm:

  • Incorporates force-field energy terms and solvation descriptors as differentiable losses during training.
  • Employs active learning: AI proposes candidates, MD simulations refine structures, and outcomes retrain the model.
  • Integrates uncertainty quantification to prioritize compounds with high predictive confidence, minimizing false positives.

Such convergence of physics and learning has cut PROTAC design cycles from months to weeks in some industrial case studies [18].

Challenges and Future Outlook

AI in PROTAC design is promising but not without hurdles:

  • Data scarcity and bias: Negative (non-degrading) examples are rare, and datasets skew toward CRBN/VHL systems.
  • Representation gaps: Many models rely on 2D graphs or insufficient handling of conformational flexibility.
  • Generalization across targets: Models trained on kinase degraders may not transfer to GPCRs or transcription factors.
  • Synthesis and ADME gap: AI may propose chemically elegant but synthetically impractical molecules.

Conclusion: Toward an Intelligent Degrader-Design Ecosystem

AI has become the silent catalyst transforming targeted protein degradation. From DeLinker to PROTAC-RL, and from PROTAC-DB 3.0 to AlphaFold-Multimer, PROTAC design is entering a data-driven era were learning complements human intuition.

The next generation of PROTAC discovery will not soley rely on brute-force synthesis, exhaustive simulations, nor on learning-guided reasoning where molecular intelligence designs with intent, precision, and speed.

As AI continues to mature, its true impact will be measured not only by faster discovery, but by smarter, safer, and more accessible therapeutics bridging chemistry, computation, and biology in ways that were unimaginable a decade ago.

At the end of the day, PROTAC design cannot be achieved in isolation or in silos. It thrives on the synergy between medicinal-chemistry intuition, informatics-driven insights, and AI-enabled design capabilities. When these three domains converge, efficient and rational PROTAC design becomes not just possible but reproducible, scalable and translatable.

Efficient PROTAC Design = Medicinal Chemistry Intuition × Informatics Insight × Data-driven Intelligence

At Centella, we believe in this integrated vision. Our platform combines the collective strengths of medicinal chemistry expertise, informatics frameworks, and real-time AI innovation to accelerate the design of next-generation degraders. If you’re exploring partnerships or seeking collaborative expertise in designing efficient, mechanism-guided PROTACs, reach out to us and let’s redefine degrader discovery together

References:

  1. Ge, J., Li, S., Weng, G., Wang, H., Fang, M., Sun, H., … & Hou, T. (2025). PROTAC-DB 3.0: an updated database of PROTACs with extended pharmacokinetic parameters. Nucleic acids research, 53(D1), D1510-D1515.
  2. https://protacpedia.weizmann.ac.il/
  3. Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., & Dahl, G. E. (2017, July). Neural message passing for quantum chemistry. In International conference on machine learning (pp. 1263–1272). Pmlr.
  4. Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., … & Pande, V. (2018). MoleculeNet: a benchmark for molecular machine learning. Chemical science, 9(2), 513–530.
  5. Evans, R., O’Neill, M., Pritzel, A., Antropova, N., Senior, A., Green, T., … & Hassabis, D. (2021). Protein complex prediction with AlphaFold-Multimer. biorxiv, 2021–10.
  6. Imrie F., Bradley A. R., van der Schaar M., Deane C. M. (2020). Deep generative models for 3D linker design. J. Chem. Inf. Model., 60(4), 1983–1995. https://doi.org/10.1021/acs.jcim.9b01120
  7. Guo, J., Knuth, F., Margreitter, C., Janet, J. P., Papadopoulos, K., Engkvist, O., & Patronov, A. (2023). Link-INVENT: generative linker design with reinforcement learning. Digital Discovery, 2(2), 392–408.
  8. Igashov, I., Stärk, H., Vignac, C. et al. Equivariant 3D-conditional diffusion model for molecular linker design. Nat Mach Intell 6, 417–427 (2024). https://doi.org/10.1038/s42256-024-00815-9
  9. Li, F., Hu, Q., Zhou, Y., Yang, H., & Bai, F. (2024). DiffPROTACs is a deep learning-based generator for proteolysis targeting chimeras. Briefings in bioinformatics, 25(5), bbae358. https://doi.org/10.1093/bib/bbae358
  10. Wurz, R. P., Rui, H., Dellamaggiore, K., Ghimire-Rijal, S., Choi, K., Smither, K., … & Vaish, A. (2023). Affinity and cooperativity modulate ternary complex formation to drive targeted protein degradation. Nature communications, 14(1), 4177.
    1. Mai, H., Zimmer, M. H., & Miller, T. F., 3rd (2023). Exploring PROTAC Cooperativity with Coarse-Grained Alchemical Methods. The journal of physical chemistry. B, 127(2), 446–455. https://doi.org/10.1021/acs.jpcb.2c05795
  11. Zheng, S., Tan, Y., Wang, Z. et al. Accelerated rational PROTAC design via deep learning and molecular simulations. Nat Mach Intell 4, 739–748 (2022). https://doi.org/10.1038/s42256-022-00527-y
  12. Erazo, F., Dunlop, N., Jalalypour, F., & Mercado, R. (2025). Enhancing PROTAC Ternary Complex Prediction with Ligand Information in AlphaFold 3.
  13. Pereira, G. P., Gouzien, C., Souza, P. C., & Martin, J. (2025). Challenges in predicting PROTAC-mediated protein–protein interfaces with AlphaFold reveal a general limitation on small interfaces. Bioinformatics Advances, 5(1), vbaf056.
  14. Li, F., Hu, Q., Zhang, X., Sun, R., Liu, Z., Wu, S., Tian, S., Ma, X., Dai, Z., Yang, X., Gao, S., & Bai, F. (2022). DeepPROTACs is a deep learning-based targeted degradation predictor for PROTACs. Nature communications, 13(1), 7133. https://doi.org/10.1038/s41467-022-34807-3
  15. Ribes, S., Nittinger, E., Tyrchan, C., & Mercado, R. (2024). Modeling PROTAC degradation activity with machine learning. Artificial Intelligence in the Life Sciences, 6, 100104.
  16. Wu, Z., Zhang, O., Wang, X., Fu, L., Zhao, H., Wang, J., … & Hou, T. (2024). Leveraging language model for advanced multiproperty molecular optimization via prompt engineering. Nature Machine Intelligence, 6(11), 1359–1369.
  17. Gharbi, Y., & Mercado, R. (2024). A comprehensive review of emerging approaches in machine learning for de novo PROTAC design. Digital Discovery.

메타데이터
post_id
aa6b1507eee8
slug
ai-approaches-in-protacs-the-new-frontier-in-targeted-protein-degradation-aa6b1507eee8
url
https://medium.com/@riyaz_76592/ai-approaches-in-protacs-the-new-frontier-in-targeted-protein-degradation-aa6b1507eee8
canonical_url
https://medium.com/@riyaz_76592/ai-approaches-in-protacs-the-new-frontier-in-targeted-protein-degradation-aa6b1507eee8
author_url
https://medium.com/@riyaz_76592
status
ok
fetched_at
2026-07-16 01:05:39