Evaluation of Large Language Model-Based Chatbots' Accuracy in Responding to CBCT-Related Questions in Endodontics: A Guideline-Based Assessment
Eurasian Dental Research. 2026;4(1):21-26.
ABSTRACT
Aim: To evaluate the accuracy and guideline compliance of ChatGPT and Claude regarding the use of cone beam computed tomography (CBCT) in endodontics, based on the European Society of Endodontology (ESE) position statement: use of CBCT in endodontics.
Material and method: A structured question set comprising 32 true/false statements and 8 open-ended questions was developed based on the ESE position statement. Questions were reviewed by an endodontic specialist and covered four main categories: CBCT indications and justification, radiation dose and imaging parameters, technical application and image interpretation, and training requirements and clinical responsibility. ChatGPT version 5.2 and Claude 4.5 Sonnet were evaluated in separate chat sessions without chat history. All responses were evaluated by a single investigator based on the ESE criteria. Data collection was completed on January 14, 2026.
Results: Both ChatGPT 5.2 and Claude 4.5 Sonnet achieved an overall accuracy of 97.5% (39/40 correct responses). For true/false statements, both chatbots correctly answered 31 out of 32 questions (96.88%), with identical errors on question 23. For open-ended questions, both achieved 100% accuracy (8/8 correct). McNemar's test could not be computed due to the absence of discordant pairs, indicating perfect agreement. Chi-square analysis showed no statistically significant difference between the two chatbots (χ² = 0.000, p = 1.000).
Conclusion: Both ChatGPT 5.2 and Claude 4.5 Sonnet demonstrated excellent accuracy and guideline compliance regarding CBCT use in endodontics based on the ESE position statement, suggesting their potential utility for guideline-based information retrieval in dental education and clinical practice.
KEYWORDS
Artificial intelligence · Cone-beam computed tomography · Endodontics · Guideline adherence · Large language models
STUDY SUMMARY
This study measured how accurately ChatGPT 5.2 and Claude 4.5 Sonnet answered questions on the use of cone beam computed tomography in endodontics, using 32 true/false statements and 8 open-ended questions derived from a European Society of Endodontology position statement. Both models reached 97.5% overall accuracy, answered the same true/false item incorrectly, and answered all open-ended questions correctly, with no statistically significant difference between them. The findings are limited to guideline-based question answering evaluated by a single investigator.