Facilitating structure-based drug discovery with an artificial intelligence
Structure-based virtual screening (VS) via molecular docking is a pivotal approach for hit identification. Many artificial intelligence (AI)-powered protein–ligand docking and scoring methods have demonstrated impressive speed and accuracy. Retrospective benchmarking studies using enrichment rate and computational efficiency on curated datasets have corroborated their potential for discovering bioactive compounds. However, determining which method suits a specific application and implementing it efficiently remains challenging. Here we present the Comprehensive VS Platform with AI Engine (CVSP-AIE) for drug discovery from compound libraries. It integrates three AI models: KarmaDock, a fast docking model that directly updates atomic coordinates; CarsiDock, an accurate docking model that predicts protein–ligand distances and reconstructs binding poses; and RTMScore, an accurate scoring model that learns residue–atom distance distributions for affinity prediction. Their hierarchical application enables dynamical balances in screening speed and accuracy. CVSP-AIE is available as an online web server (https://cadd.zju.edu.cn/cvsp/) and a local software package. Users can efficiently initiate drug screening by uploading a protein and a known binder that defines the binding pocket. The following workflow involves (1) preprocessing, including protein structure repair and molecule standardization, (2) binding pose and affinity prediction powered by KarmaDock, CarsiDock and RTMScore and (3) postprocessing, comprising protein–ligand interaction calculation and visualization. It takes 30–45 min to hierarchically screen 100,000 compounds, and the output is a ranked list of molecules with predicted binding scores, intermolecular interaction profiles and interactive chemical space analysis. Users can also install locally the hierarchical screening module through command-line package for arbitrary-scale screening.
Comprehensive Virtual Screening (VS) Platform with Artificial Intelligence Engine (CVSP-AIE) is a tool for structure-based VS. Its online web server allows users to efficiently initiate drug screening by simply uploading target proteins and reference ligands through the web interface. The VS process is primarily based on three artificial intelligence models.
The local software package of CVSP-AIE integrates the core hierarchical VS strategy, enabling users to deploy it locally and perform large-scale VS tasks via command-line interface.
This is a preview of subscription content, access via your institution
Prices may be subject to local taxes which are calculated during checkout
All datasets that were analyzed within this Protocol can be obtained from https://huggingface.co/gushukai/HierVS.
All source codes of this protocol are available for use under MIT license and are available via https://huggingface.co/gushukai/HierVS, via Zenodo at https://zenodo.org/records/18073542 (ref. 50) and via GitHub https://github.com/shukai1997/HierVS (ref. 51). The CVSP-AIE web platform is freely available at https://cadd.zju.edu.cn/cvsp.
Berdigaliyev, N. & Aljofan, M. An overview of drug discovery and development. Future Med. Chem. 12, 939–947 (2020).
Article CAS PubMed Google Scholar
Hughes, J. P., Rees, S., Kalindjian, S. B. & Philpott, K. L. Principles of early drug discovery. Br. J. Pharmacol. 162, 1239–1249 (2011).
Article CAS PubMed PubMed Central Google Scholar
Jones, A. & Clifford, L. Drug discovery alliances. Nat. Rev. Drug Discov. 4, 807–808 (2005).
Zhu, T. et al. Hit identification and optimization in virtual screening: practical recommendations based on a critical literature analysis. J. Med. Chem. 56, 6560–6572 (2013).
Bleicher, K. H., Bohm, H. J., Muller, K. & Alanine, A. I. Hit and lead generation: beyond high-throughput screening. Nat. Rev. Drug Discov. 2, 369–378 (2003).
Inglese, J. et al. High-throughput screening assays for the identification of chemical probes. Nat. Chem. Biol. 3, 466–479 (2007).
Bajorath, J. Integration of virtual and high-throughput screening. Nat. Rev. Drug Discov. 1, 882–894 (2002).
Schneider, G. Virtual screening: an endless staircase? Nat. Rev. Drug Discov. 9, 273–276 (2010).
Lyu, J., Irwin, J. J. & Shoichet, B. K. Modeling the expansion of virtual screening libraries. Nat. Chem. Biol. 19, 712–718 (2023).
Deane, C. & Mokaya, M. A virtual drug-screening approach to conquer huge chemical libraries. Nature 601, 322–323 (2022).
Shoichet, B. K. Virtual screening of chemical libraries. Nature 432, 862–865 (2004).
Ripphausen, P., Nisius, B. & Bajorath, J. State-of-the-art in ligand-based virtual screening. Drug Discov. Today 16, 372–376 (2011).
Lionta, E., Spyrou, G., Vassilatis, D. K. & Cournia, Z. Structure-based virtual screening for drug discovery: principles, applications and recent advances. Curr. Top Med. Chem. 14, 1923–1938 (2014).
Lyne, P. D. Structure-based virtual screening: an overview. Drug Discov. Today 7, 1047–1055 (2002).
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).
Burley, S. K. et al. RCSB Protein Data Bank (RCSB.org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res. 51, D488–D508 (2023).
Gu, S. et al. Evaluation of AlphaFold2 structures for hit identification across multiple scenarios. J. Chem. Inf. Model 64, 3630–3639 (2024).
Shen, C. et al. Boosting protein–ligand binding pose prediction and virtual screening based on residue-atom distance likelihood potential and graph transformer. J. Med. Chem. 65, 10691–10706 (2022).
Zhang, X. et al. Efficient and accurate large library ligand docking with KarmaDock. Nat. Comput. Sci. 3, 789–804 (2023).
Cai, H. et al. CarsiDock: a deep learning paradigm for accurate protein–ligand docking and screening based on large-scale pre-training. Chem. Sci. 15, 1449–1471 (2024).
Gu, S. et al. Benchmarking AI-powered docking methods from the perspective of virtual screening. Nat. Mach. Intell. 7, 509–520 (2025).
Corso, G., Stärk, H., Jing, B., Barzilay, R. & Jaakkola, T. Diffdock: diffusion steps, twists, and turns for molecular docking. In International Conference on Learning Representations (ICLR, 2023).
Li, Y., Li, L., Wang, S. & Tang, X. EQUIBIND: a geometric deep learning-based protein–ligand binding prediction method. Drug Discov. Ther 17, 363–364 (2023).
Lu, W. et al. TANKBind: trigonometry-aware neural networks for drug-protein binding structure prediction. Adv. Neural Inf. Process. Syst. 35, 7236–7249 (2022).
Friesner, R. A. et al. Glide: a new approach for rapid, accurate docking and scoring. 1. Method and assessment of docking accuracy. J. Med. Chem. 47, 1739–1749 (2004).
Wang, Z. et al. Comprehensive evaluation of ten docking programs on a diverse set of protein–ligand complexes: the prediction accuracy of sampling power and scoring power. Phys. Chem. Chem. Phys. 18, 12964–12975 (2016).
Ruiz-Carmona, S. et al. rDock: a fast, versatile and open source program for docking ligands to proteins and nucleic acids. PLoS Comput. Biol. 10, e1003571 (2014).
Article PubMed PubMed Central Google Scholar
Jain, A. N. Surflex: fully automatic flexible molecular docking using a molecular similarity-based search engine. J. Med. Chem. 46, 499–511 (2003).
Xia, S., Chen, E. & Zhang, Y. Integrated molecular modeling and machine learning for drug design. J. Chem. Theory Comput. 19, 7478–7495 (2023).
Wishart, D. S. et al. DrugBank: a comprehensive resource for in silico drug discovery and exploration. Nucleic Acids Res. 34, D668–D672 (2006).
Chi, X. et al. Discovery and characterization of novel FAK inhibitors for breast cancer therapy via hybrid virtual screening, biological evaluation and molecular dynamics simulations. Bioorg. Chem. 159, 108400 (2025).
Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. Preprint at bioRxiv https://doi.org/10.1101/2025.06.14.659707 (2025).
Team, B. A. A. S. et al. Protenix—advancing structure prediction through a comprehensive AlphaFold3 reproduction. Preprint at bioRxiv https://doi.org/10.1101/2025.01.08.631967 (2025).
Bauer, M. R., Ibrahim, T. M., Vogel, S. M. & Boeckler, F. M. Evaluation and optimization of virtual screening workflows with DEKOIS 2.0—a public library of challenging docking benchmark sets. J. Chem. Inf. Model. 53, 1447–1462 (2013).
Chen, L. et al. Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening. PLoS ONE 14, e0220113 (2019).
Chatterjee, A. et al. Improving the generalizability of protein–ligand binding predictions with AI-Bind. Nat. Commun. 14, 1989 (2023).
Mohammad, T., Mathur, Y. & Hassan, M. I. InstaDock: a single-click graphical user interface for molecular docking-based virtual high-throughput screening. Brief. Bioinform. 22, https://doi.org/10.1093/bib/bbaa279 (2021).
Mo, Q., Xu, Z., Yan, H., Chen, P. & Lu, Y. VSTH: a user-friendly web server for structure-based virtual screening on Tianhe-2. Bioinformatics https://doi.org/10.1093/bioinformatics/btac740 (2023).
Gan, J. H. et al. DrugRep: an automatic virtual screening server for drug repurposing. Acta Pharmacol. Sin. 44, 888–896 (2023).
Guedes, I. A. et al. DockThor-VS: a free platform for receptor-ligand virtual screening. J. Mol. Biol. 436, 168548 (2024).
Zhou, G. et al. An artificial intelligence accelerated virtual screening platform for drug discovery. Nat. Commun. 15, 7761 (2024).
Gao, B. et al. DrugCLIP: contrastive protein-molecule representation learning for virtual screening. Adv Neural Inf. Process. Syst. 36, 44595–44614 (2024).
Buttenschoen, M., Morris, G. M. & Deane, C. M. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chem. Sci. 15, 3130–3139 (2024).
Masters, M. R., Mahmoud, A. H. & Lill, M. A. Investigating whether deep learning models for co-folding learn the physics of protein–ligand interactions. Nat. Commun. 16, 8854 (2025).
Li, Y. et al. Decoding the limits of deep learning in molecular docking for drug discovery. Chem. Sci. 16, 17374–17390 (2025).
Landrum, G. et al. RDKit: open-source cheminformatics. GitHub and SourceForge https://www.rdkit.org/ (2025).
Tran-Nguyen, V. K., Jacquemard, C. & Rognan, D. LIT-PCBA: an unbiased data set for machine learning and virtual screening. J. Chem. Inf. Model. 60, 4263–4273 (2020).
Sastry, G. M., Adzhigirey, M., Day, T., Annabhimoju, R. & Sherman, W. Protein and ligand preparation: parameters, protocols, and influence on virtual screening enrichments. J. Comput. Aided Mol. Des. 27, 221–234 (2013).
Gu, S. The Docker image for ‘CVSP-AIE: a comprehensive virtual screening platform with artificial intelligence engine’. Zenodo https://zenodo.org/records/18073542 (2025).
Gu, S. Code for ‘Facilitating structure-based drug discovery with an artificial intelligence driven virtual screening platform’. Zenodo https://zenodo.org/records/19448262 (2026).
Chen, L. et al. TransformerCPI: improving compound-protein interaction prediction by sequence-based deep learning with self-attention mechanism and label reversal experiments. Bioinformatics 36, 4406–4414 (2020).
Eberhardt, J., Santos-Martins, D., Tillack, A. F. & Forli, S. AutoDock Vina 1.2.0: new docking methods, expanded force field, and python bindings. J. Chem. Inf. Model. 61, 3891–3898 (2021).
Verdonk, M. L., Cole, J. C., Hartshorn, M. J., Murray, C. W. & Taylor, R. D. Improved protein–ligand docking using GOLD. Proteins 52, 609–623 (2003).
Santos-Martins, D. et al. Accelerating AutoDock4 with GPUs and gradient-based local search. J. Chem. Theory Comput. 17, 1060–1073 (2021).
Fey, M. & Lenssen, J. E. Fast graph representation learning with PyTorch Geometric. Preprint at https://arxiv.org/abs/1903.02428 (2019).
Paszke, A. et al. Pytorch: an imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32 (eds Wallach, H. M. et al.) 8026–8037 (ACM, 2019).
This work was financially supported by the National Key Research and Development Program of China (grant no. 2024YFA1307501), the National Natural Science Foundation of China (grant nos. 22220102001, 92370130, 22503081, 82473843 and 82204279), the Youth Project of Hunan Administration of Traditional Chinese Medicine (grant no. 2021185), China Postdoctoral Science Foundation (grant no. 2024M762886) and Postdoctoral Fellowship Program of CPSF (grant no. GZC20252381).
These authors contributed equally: Shukai Gu, Xujun Zhang, Mengwu Xiao.
College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, China
Shukai Gu, Xujun Zhang, Yuntao Qian, Hao Luo, Hongyan Du, Odin Zhang, Minjie Mou, Tingting Fu, Xiaorui Wang, Jingxuan Ge, Feng Zhu, Tingjun Hou & Yu Kang
Faculty of Applied Science, Macao Polytechnic University, Macao, China
Shukai Gu, Bo Liu, Xiaojun Yao & Huanxiang Liu
School of Pharmacy, Hunan University of Chinese Medicine, Changsha, China
Department of Clinical Pharmacy, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China
Search author on:PubMed Google Scholar
S.G. and Y.K. conceived the idea and designed the entire research. S.G., X.Z., M.X. and Y.Q. wrote codes. S.G., B.L., H.Luo, M.M. and H.D. conducted web server test. O.Z., T.F., X.W., J.G., C.S., S.Z., G.W., J.Z., H.S., M.M., J.W., F.Z. and X.Y. finished result analysis. S.G., H.Liu, T.H. and Y.K. visualized the results. S.G., H.Liu, T.H. and Y.K. wrote the manuscript.
Correspondence to Huanxiang Liu, Tingjun Hou or Yu Kang.
The authors declare no competing interests.
Nature Protocols thanks Sandro Cosconati, Vincent Zoete and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Gu, S. et al. Nat. Mach. Intell. 7, 509–520 (2025): https://doi.org/10.1038/s42256-025-00993-0
Zhang, X. et al. Nat Comput Sci 3, 789–804 (2023): https://doi.org/10.1038/s43588-023-00511-5
Cai, H. et al. Chem. Sci. 15, 1449–1471 (2024): https://doi.org/10.1039/D3SC05552C
Shen, C. et al. J. Med. Chem. 65, 10691–10706 (2022): https://doi.org/10.1021/acs.jmedchem.2c00991
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Gu, S., Zhang, X., Xiao, M. et al. Facilitating structure-based drug discovery with an artificial intelligence-driven virtual screening platform. Nat Protoc (2026). https://doi.org/10.1038/s41596-026-01389-z
Version of record: 24 June 2026
DOI: https://doi.org/10.1038/s41596-026-01389-z
Related Stories
AI News
Reality Check: A Guide to Navigating the AI Information Flood
11 minutes ago
AI News
Protecting Digital Privacy In The Artificial Intelligence Era
11 minutes ago
AI News
A third of web pages published since ChatGPT launched were written by AI, study finds
1 hour ago
AI News
Stripe didn’t really buy OpenRouter because of the ‘singularity’
1 hour ago
AI News
Etched’s valuation doubles to $21B in a month
2 hours ago
AI News
China is winning one AI race, the US another
2 hours ago
AI News
What about the ethical challenges facing major AI developers?
3 hours ago
AI News
Seeking Public Comment! Using Artificial Intelligence for Cybersecurity Framework 2.0 Analysis and Reporting
4 hours ago