JHSM

Journal of Health Sciences and Medicine (JHSM) is an unbiased, peer-reviewed, and open access international medical journal. The Journal publishes interesting clinical and experimental research conducted in all fields of medicine, interesting case reports, and clinical images, invited reviews, editorials, letters, comments, and related knowledge.

EndNote Style
Index
Original Article
Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam
Aims: The aim of this study was to compare the accuracy of ChatGPT (GPT-5), Gemini (Gemini 2.5 Pro), and Grok (SuperGrok) on Dental Specialty Exam (DUS) endodontics questions and to examine differences by exam year and item format (text-based vs figure-based). Methods: Endodontics questions from Turkish Assessment, Selection and Placement Center (ÖSYM) DUS exams (2012-2021) were reviewed. Of 130 questions, three items officially canceled by ÖSYM were excluded, leaving 127 questions (122 text-based, 5 figure-based). Figure-based items included periapical radiographs, clinical photographs, or schematic diagrams. Questions were submitted to each model using default settings with one response per item; no tuning or repeat runs were performed. Responses were scored as correct/incorrect using the official answer key. Fisher’s exact test and, when required, Monte Carlo-corrected Fisher’s exact test were applied (p


1. Turosz N, Checinska K, Checinski M, Brzozowska A, Nowak Z, Sikora M. Applications of Artificial Intelligence in the analysis of dental panoramic radiographs: an overview of systematic reviews. Dentomaxillofac Radiol. 2023;52(7):20230284. doi:10.1259/dmfr.20230284
2. Schwendicke F, Samek W, Krois J. Artificial Intelligence in dentistry: opportunities and challenges. J Dent Res. 2020;99(7):769-774. doi:10. 1177/0022034520915714
3. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/s41591-023-02448-8
4. Alhaidry HM, Fatani B, Alrayes JO, Almana AM, Alfhaed NK. ChatGPT in dentistry: a comprehensive review. Cureus. 2023;15(4):e38317. doi:10. 7759/cureus.38317
5. Riedemann L, Labonne M, Gilbert S. The path forward for large language models in medicine is open. Npj Digit Med. 2024;7(1):339. doi: 10.1038/s41746-024-01344-w
6. Dave M, Tattar R, Alafaleg R, et al. Performance of large language models (ChatGPT4-0, Grok2 and Gemini) in UK dentistry and dental hygiene and therapy assessments. Br Dent J. 2025. doi:10.1038/s41415-025-8383-2
7. Zhou J, Li H, Chen S, et al. Large language models in biomedicine and healthcare. npj Artif Intell. 2025;1:44. doi:10.1038/s44387-025-00047-1
8. Eggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent. 2023;35(7):1098-1102. doi:10.1111/jerd.13046
9. Karobari MI, Adil AH, Basheer SN, et al. Evaluation of the diagnostic and prognostic accuracy of Artificial Intelligence in endodontic dentistry: a comprehensive review of literature. Comput Math Methods Med. 2023;2023:7049360. doi:10.1155/2023/7049360
10. Suárez A, Díaz-Flores García V, Algar J, Gómez Sánchez M, Llorente de Pedro M, Freire Y. Evaluating the consistency and accuracy of endodontic question answers generated by ChatGPT. Int Endod J. 2024; 57(1):108-113. doi:10.1111/iej.13985
11. Fontenele RC, Jacobs R. Unveiling the power of Artificial Intelligence for image-based diagnosis and treatment in endodontics: an ally or adversary? Int Endod J. 2025;58(2):155-170. doi:10.1111/iej.14163
12. Pul U, Schwendicke F. Artificial Intelligence for detecting periapical radiolucencies: a systematic review and meta-analysis. J Dent. 2024;147: 105104. doi:10.1016/j.jdent.2024.105104
13. Jalali P, Mohammad-Rahimi H, Wang FM, et al. Performance of 7 Artificial Intelligence chatbots on board-style endodontic questions. J Endod. 2025;51(10):1413-1419. doi:10.1016/j.joen.2025.06.014
14. Durmazpinar PM, Ekmekci E. Comparing diagnostic skills in endodontic cases: dental students versus ChatGPT-4o. BMC Oral Health. 2025;25:457. doi:10.1186/s12903-025-05857-y
15. ÖSYM. Past questions from the Specialization Education Entrance Exam in Dentistry (DUS) [Internet]. Ankara: Measurement, Selection, and Placement Center Presidency; [cited 2026 Jan 24]. Available from: https://www.osym.gov.tr/TR,15070/dus-cikmis-sorular.html
16. Dashti M, Takabi M, Rezaee J. Performance of ChatGPT 3.5 and 4 on U.S. dental examinations: the INBDE, ADAT, and DAT. Imaging Sci Dent. 2024;54(3):271-275. doi:10.5624/isd.20240037
17. Nguyen HC, Dang HP, Nguyen TL, Hoang V, Nguyen VA. Accuracy of latest large language models in answering multiple choice questions in dentistry: a comparative study. PLoS One. 2025;20(1):e0317423. doi:10. 1371/journal.pone.0317423
18. Jeong H, Han SS, Yu Y, Kim S, Jeon KJ. How well do large language model-based chatbots perform in oral and maxillofacial radiology? Dentomaxillofac Radiol. 2024;53(6):390-395. doi:10.1093/dmfr/twae021
19. Bilgin Avşar D, Ertan AA. A comparative study of ChatGPT-3.5 and Gemini’s performance in answering prosthetic dentistry questions in the Dentistry Specialty Exam: cross-sectional study. Turkiye Klinikleri J Dental Sci. 2024;30(4):668-673. doi:10.5336/dentalsci.2024-104610
20. Meriç E. Comparative analysis of ChatGPT-4o Plus and Gemini in answering pediatric dentistry questions for the Dentistry Specialization Exam: cross-sectional study. Turkiye Klinikleri J Dental Sci. 2025;31(4): 555-561. doi:10.5336/dentalsci.2025-108696
21. Peker RB. Performance of ChatGPT, ChatGPT Plus, Gemini, and Microsoft Copilot large language models on oral and maxillofacial radiology questions in the Turkish dentistry specialization exam. Yeditepe Dental Journal. 2025;21(3):130-135. doi:10.5505/yeditepe.2025. 93265
22. Tosun B, Yılmaz ZS. Comparison of Artificial Intelligence systems in answering prosthodontics questions from the dental specialty exam in Turkey. J Dent Sci. 2025;20(3):1454-1459. doi:10.1016/j.jds.2025.01.025
23. Yılmaz BE, Gökkurt Yılmaz BN, Özbey F. Artificial Intelligence performance in answering multiple-choice oral pathology questions: a comparative analysis. BMC Oral Health. 2025;25:573. doi:10.1186/s12903-025-05926-2
24. Özden, S, Erener, H. Comparative Evaluation of ChatGPT-4o, Gemini 2.5 Pro and Grok-4 in Answering Orthodontics Questions from the Dentistry Specialty Examination. Lokman Hekim Health Sci. 2026;6(1): 126-133. doi: 10.14744/lhhs.2025.42941
25. Başkan HK, Başkan B. Performance comparison of large language models on pediatric dentistry questions in the Turkish dentistry specialization examination. BMC Med Educ. 2025;25:1734. doi:10.1186/s12909-025-08315-z
26. Arılı Öztürk E, Turan Gökduman C, Çanakçi BC. Evaluation of the performance of ChatGPT-4 and ChatGPT-4o as a learning tool in endodontics. Int Endod J. 2025. doi:10.1111/iej.14217
27. Çekiç EC, Tavşan O. Evaluating large language models using national endodontic specialty examination questions: are they ready for real-world dentistry? BMC Med Educ. 2025;25:1308. doi:10.1186/s12909-025-07896-z
28. Gürsu Şahin E. Retrospective analysis of endodontics questions asked in the specialty examination in dentistry. Turkiye Klinikleri J Dental Sci. 2024;30(1):107-114. doi:10.5336/dentalsci.2023-100386
29. Roustan D, Bastardot F. The clinicians’ guide to large language models: a general perspective with a focus on hallucinations. Interact J Med Res. 2025;14:e59823. doi:10.2196/59823
Volume 9, Issue 3, 2026
Page : 767-773
_Footer