Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models
For patients and families
In plain language
An automatic summary of structured registry data. It is an orientation aid, not a substitute for the official protocol or a physician assessment.
- What is being studied
- The protocol lists: ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, Claude Opus 4.7.
- Who it may be relevant to
- Registry conditions: Pulp and Periapical Tissue Disease. Basic parameters: from 16 years · All.
- What needs checking
- Age, condition and sex are only basic indicators. Prior treatment, laboratory values and other mandatory requirements appear in the eligibility criteria below.
- Where it takes place
- Center list to be confirmed — check the primary protocol.
- Next step
- Save the trial, show it to the treating physician, and confirm current recruitment with the study center. Costs, documents and travel →
Unsure about the terms? Read our patient guide →
Official title
Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard
Overview
This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.
Detailed description
This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.
Interventions
- Device ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, Claude Opus 4.7
Large language model used to analyze standardized clinical information and periapical radiographs to provide pulpal and periapical diagnosis and endodontic case difficulty assessment according to AAE criteria
Primary outcome measures
- Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis [Time frame: At baseline]
Secondary outcome measures (1)
- Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment [Time frame: At baseline]
Eligibility criteria
Inclusion criteria
- Age above 16 years old.
- Systemically healthy patient (ASA I or II).
- Requiring endodontic treatment or retreatment
- Patient's acceptance to participate in the study
Exclusion criteria
- Medically compromised patients.
- Pregnant women.
- Traumatic dental injuries
- Low quality periapical radiograph
Criteria are shown verbatim from the registry (in English). Final eligibility is always assessed by the study center.
Healthy volunteers: Yes
Study design
- Observational model
- Other
Study locations
Center list to be confirmed — check the primary protocol.
Identifiers
NCT: NCT07732985 · New ENDO7.1.1