Introduction. Large Language Models (LLMs), artificial intelligence–based programs trained on arrays of text data and capable of generating coherent text responses to arbitrary queries, have revolutionized information interaction in many fields over the past two years, including medicine. However, urology poses a special challenge for LLMs for several reasons: highly specialized terminology, a variety of clinical scenarios, the need to take into account current clinical recommendations, and the high importance of accurate numerical data.
Objective. To evaluate the clinical accuracy and safety of 18 LLMs in answering urological questions and compare them with a retrieval-augmented generation (RAG) system based on a verified evidence base.
Materials and methods. A prospective comparative study included two directions: factual accuracy assessment of 17 LLMs on 28 specific urological questions with verified ground truth (115 facts, 2% tolerance); clinical evaluation of 18 LLMs on 120 clinical cases from real consultations com-pared with reference urologist answers. Models were stratified into 4 tiers: flagship (Tier 1, n=7), strong (Tier 2, n=4), open-source (Tier 3, n=4), and small (Tier 4, n=3). RAG system: Qwen3-80B + ChromaDB (2.23M chunks, 17.947 documents, 19 domains). Quality assessment: LLM-as-judge (Gemini 2.5 Flash), 2.260 clinical evaluations. Statistics: ANOVA, Cohen's d, 95% confidence intervals.
Results. RAG outperformed all bare LLMs in accuracy: 4.83 vs 4.09 (best model, Claude Opus 4.6), Cohen's d = 1.05 (p<0.001). Differences between Tier 1, Tier 2, and Tier 3 were negligible (Cohen's d = 0.07–0.15), while Tier 4 demonstrated clinically unacceptable accuracy (2.52/5). The greatest RAG advantage was observed in oncourology (Δ=+1.47–1.89) and surgical techniques (Δ=+1.77). All bare LLMs demonstrated substantial safety issues (0.85–3.09 per case), while empathy remained consistently high (4.18–4.85), creating a false impression of competence.
Conclusion. Bare LLMs of all tiers demonstrate a high hallucination rate in clinical urology. The RAG system eliminates this limitation with a large effect size. Flagship model status does not guarantee clinical accuracy.
