Quality assessment of artificial intelligence responses in erectile dysfunction: a comparative study based on EAU recommendations
Artificial intelligence (AI)–based language models are increasingly explored as tools for interpreting and applying clinical guideline recommendations. In urology, the European Association of Urology (EAU) recently introduced a guideline-specific chatbot; however, its comparative performance relative to contemporary general-purpose large language models (LLMs) remains unclear. In this structured comparative study, five AI systems—the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot – Smart GPT-5, and Perplexity Pro—were evaluated using 13 clinical questions derived directly from strongly recommended statements in the EAU erectile dysfunction (ED) guidelines. Responses were independently assessed by three senior reviewers across five predefined domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome of the study was the composite performance score, which was calculated as the mean of the five domain scores. Inter-rater reliability was calculated using ICC(2,k), and differences among models were analyzed with the Friedman test followed by Holm-adjusted Wilcoxon post-hoc comparisons. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40–4.73)] and the EAU Guidelines Bot [4.53 (4.47–4.80)], followed by ChatGPT-5 [4.27 (4.07–4.47)]. Lower composite scores were observed for Copilot – Smart GPT-5 [3.73 (3.40–3.87)] and Perplexity Pro [3.60 (3.47–3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. These findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured ED scenarios. However, variability across domains—particularly in structure and clinical utility—and modest differences in composite performance suggest that these models should be interpreted as supportive tools rather than definitive clinical decision-making systems, requiring further validation in real-world settings.
This is a preview of subscription content, access via your institution
Prices may be subject to local taxes which are calculated during checkout
Data are available from the corresponding author on reasonable request.
Impotence: NIH Consensus Development Panel on Impotence. JAMA. 1993;270:83–90. https://doi.org/10.1001/JAMA.1993.03510010089036.
McKinlay JB. The worldwide prevalence and epidemiology of erectile dysfunction. Int J Impot Res. 2000;12:S6–11. https://doi.org/10.1038/SJ.IJIR.3900567.
Salonia, Capogrosso A, Boeri P, Cocci L, Corona G A, Dinkelman-Smit M, et al. European Association of Urology guidelines on male sexual and reproductive health: 2025 update on male hypogonadism, erectile dysfunction, premature ejaculation, and Peyronie’s disease. Eur Urol. 2025;88:76–102. https://doi.org/10.1016/J.EURURO.2025.04.010.
Hussain W, Khoriba G, Maity S, Jyoti Saikia M. Large language models in healthcare and medical applications: a review. Bioengineering. 2025;12:631. https://doi.org/10.3390/BIOENGINEERING12060631.
Baturu M, Solakhan M, Kazaz TG, Bayrak O. Frequently asked questions on erectile dysfunction: evaluating artificial intelligence answers with expert mentorship. Int J Impot Res. 2024;37:310–4. https://doi.org/10.1038/s41443-024-00898-3.
EAU Guidelines Bot. Guideline Bot - Uroweb 2025. https://uroweb.org/chat//6704965 (accessed September 15, 2025).
Razdan S, Siegal AR, Brewer Y, Sljivich M, Valenzuela RJ. Assessing ChatGPT’s ability to answer questions pertaining to erectile dysfunction: can our patients trust it?. Int J Impot Res. 2023;36:734–40. https://doi.org/10.1038/s41443-023-00797-z.
Şahin MF, Ateş H, Keleş A, Özcan R, Doğan Ç, Akgül M, et al. Responses of five different artificial intelligence chatbots to the top searched queries about erectile dysfunction: a comparative Analysis. J Med Syst. 2024;48:38. https://doi.org/10.1007/S10916-024-02056-0.
Article PubMed PubMed Central Google Scholar
ChatGPT version 5.0 [Internet]. OpenAI [cited 2025 Sep 15]. Available from: https://chatgpt.com/
Gemini 2.5 Pro [Internet]. Google DeepMind [cited 2025 Sep 15]. Available from: https://gemini.google.com/
Copilot (GPT-5-based) [Internet]. GitHub [cited 2025 Sep 15]. Available from: https://copilot.microsoft.com/
Perplexity Pro [Internet]. Perplexity AI [cited 2025 Sep 15]. Available from: https://www.perplexity.ai/
Joshi A, Kale S, Chandel S, Pal D. Likert scale: explored and explained. Br J Appl Sci Technol. 2015;7:396–403. https://doi.org/10.9734/BJAST/2015/14975.
Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15:155–63. https://doi.org/10.1016/J.JCM.2016.02.012.
Caglar U, Yildiz O, Meric A, Ayranci A, Gelmis M, Sarilar O, et al. Evaluating the performance of ChatGPT in answering questions related to pediatric urology. J Pediatr Urol. 2024;20:26.e1–26.e5. https://doi.org/10.1016/j.jpurol.2023.08.003.
Zhu L, Mou W, Chen R. Can the ChatGPT and other large language models with internet-connected database solve the questions and concerns of patient with prostate cancer and help democratize medical knowledge? J Transl Med. 2023;21:269. https://doi.org/10.1186/s12967-023-04123-5.
Malak A, Şahin MF. How useful are current chatbots regarding urology patient information? Comparison of the ten most popular chatbots’ responses about female urinary incontinence. J Med Syst. 2024;48:102. https://doi.org/10.1007/s10916-024-02125-4.
May M, Körner-Riffard K, Kollitsch L, Burger M, Brookman-May SD, Rauchenwald M, et al. Evaluating the efficacy of AI chatbots as tutors in urology: a comparative analysis of responses to the 2022 in-service assessment of the European Board of Urology. Urol Int. 2024;108:359–66. https://doi.org/10.1159/000537854.
Karches KE. Against the iDoctor: why artificial intelligence should not replace physician judgment. Theor Med Bioeth. 2018;39:91–110. https://doi.org/10.1007/s11017-018-9442-3.
Lebovitz S, Lifshitz-Assaf H, Levina N. To engage or not to engage with AI for critical judgments: how professionals deal with opacity when using AI for medical diagnosis. Organ Sci. 2022;33:126–48. https://doi.org/10.1287/ORSC.2021.1549.
Almada M, Petit N. The EU AI Act: between the rock of product safety and the hard place of fundamental rights. Common Mark Law Rev. 2025;62:85–120. https://doi.org/10.54648/cola2025004.
The authors thank all participants who contributed to this study.
The authors received no financial support for the research, authorship, and/or publication of this article.
Department of Urology, Basaksehir Cam and Sakura City Hospital, Istanbul, Turkey
Department of Histology and Embryology, Hamidiye Institute of Health Sciences, University of Health Sciences, Istanbul, Turkey
Department of Urology, University Hospital Center of São João, Alameda Professor Hernâni Monteiro, Porto, Portugal
Urology Section, Department of Surgery, University of Catania, Catania, Italy
Urology Clinic, Città della Salute e della Scienza, University of Turin, Turin, Italy
Department of Urology, Hospital Universitario HM Sanchinarro, Instituto Investigación Sanitaria HM Hospitales and ROC Clinic, Madrid, Spain
Esther García Rojo & Esther García Rojo
Urology Unit, Department of Woman, Child and General and Specialized Surgery, University of Campania “Luigi Vanvitelli”, Naples, Italy
Department of Urology, Cleveland Clinic Abu Dhabi, Abu Dhabi, United Arab Emirates
Search author on:PubMed Google Scholar
GÇ conceptualized and designed the study. GÇ and AM performed the primary data analysis and drafted the manuscript. OA conducted the statistical analyses. GIR and MF contributed to data acquisition and provided clinical input. EGR, CM, OA, and RR revised the manuscript for important intellectual content. All authors reviewed and approved the final version of the manuscript.
Correspondence to Gökhan Çeker.
Marco Falcone is a Deputy Editor of the International Journal of Impotence Research. Afonso Morgado is an Associate Editor of the International Journal of Impotence Research. The other authors declare no competing interests.
Ethical approval was not required for this study as it did not involve human participants, animals, or identifiable patient data.
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Çeker, G., Morgado, A., Russo, G.I. et al. Quality assessment of artificial intelligence responses in erectile dysfunction: a comparative study based on EAU recommendations. Int J Impot Res (2026). https://doi.org/10.1038/s41443-026-01339-z
Version of record: 07 August 2026
DOI: https://doi.org/10.1038/s41443-026-01339-z
Related Stories
AI News
Cardinal Intelligence aims to make U of L an AI
8 minutes ago
AI News
The CEO of the unicorn AI accounting startup, Rillet, says AI is changing who consulting firms hire
38 minutes ago
AI News
7 ways AI can be used to enhance security operations
1 hour ago
AI News
Argentina Artificial Intelligence in Packaging
1 hour ago
AI News
Dr. Dre Says Artificial Intelligence Is Not a Threat to Creative Musicians
1 hour ago
AI News
Claude is down
1 hour ago
AI News
Omega Seiki Mobility Partners with Electra AI for Battery Health Intelligence
2 hours ago
AI News
Artificial intelligence in action: how AI agents are transforming businesses, public administrations and strategic sectors
2 hours ago