Меню
Идёт набор NCT07632859

Diagnostic Accuracy of GPT-4o and Claude 4.6 Sonnet in Turkish ED Anamnesis Notes

Наблюдательное Emergency Medicine Diagnostic Errors Artificial Intelligence (AI) in Diagnosis

Ориентир для пациента и семьи

Простыми словами

Автоматическая сводка по структурированным данным реестра. Она помогает сориентироваться, но не заменяет официальный протокол или оценку врача.

Что изучают
Это наблюдательное исследование: исследуемое лечение участникам по протоколу не назначают.
Кому может быть актуально
Состояния в реестре: Emergency Medicine, Diagnostic Errors, Artificial Intelligence (AI) in Diagnosis. Базовые параметры: от 18 лет · Все.
Что важно проверить
Возраст, диагноз и пол — только базовые ориентиры. Предыдущее лечение, анализы и другие обязательные условия указаны ниже в критериях участия.
Где проводится
Turkey (Türkiye)
Следующий шаг
Сохраните исследование, покажите его лечащему врачу и уточните актуальный статус у исследовательского центра. Расходы, документы и поездка →
Официальное название

Diagnostic Accuracy of Large Language Models From Emergency Department Anamnesis Notes: A Comparison of GPT-4o and Claude 4.6 Sonnet With Emergency Medicine Specialists

Обзор

This retrospective diagnostic accuracy study evaluates the ability of two large language models (LLMs) - GPT-4o (gpt-4o-2024-11-20; OpenAI) and Claude 4.6 Sonnet (claude-sonnet-4-6; Anthropic) - to generate correct diagnoses from anonymized Turkish-language emergency department (ED) anamnesis notes, and compares their performance with the diagnosis entered by the treating emergency physician. A consensus gold standard is established by three independent board-certified emergency medicine specialists who blindly review each note and vote on the primary diagnosis using ICD-10 three-character codes; the majority vote (at least 2 of 3 specialists agreeing) constitutes the reference standard. Both LLMs are evaluated using a standardized zero-shot direct prompting strategy (temperature=0, stateless API sessions). The primary outcome is diagnostic accuracy (proportion of ICD-10 chapter-level matches) and Cohen's kappa for each LLM against the gold standard. Secondary outcomes include top-3 accuracy, treating physician accuracy, inter-model agreement, and subgroup analyses by ESI triage level and ICD-10 chapter. Inter-rater reliability among the three specialists is quantified using Fleiss' kappa. Analyses are performed in Jamovi. This study represents the first evaluation of LLM diagnostic accuracy using Turkish-language clinical notes and the first to benchmark LLM performance against an independent three-specialist majority-vote gold standard rather than against the treating physician's own diagnosis.

Подробное описание

STUDY DESIGN: Retrospective diagnostic accuracy study, STARD-AI 2025 reporting, single center, cohort design.

AI INDEX TESTS: (1) GPT-4o (model version gpt-4o-2024-11-20; OpenAI API). (2) Claude 4.6 Sonnet (model version claude-sonnet-4-6; Anthropic API). Both accessed via Python (Google Colab). Temperature=0 for reproducibility. Zero-shot, stateless sessions - no cross-case context. No task-specific fine-tuning or additional training applied; models used as-is via API.

MODEL INTERPRETABILITY: Model interpretability analyses (such as SHAP, Grad-CAM, or layer-attribute visualizations) are not applicable to this study. Because GPT-4o and Claude 4.6 Sonnet are accessed as black-box models through proprietary, closed-source commercial APIs, internal model weights, gradients, and attention architectures are structurally inaccessible for post-hoc interpretability computations.

REFERENCE STANDARD: Three board-certified emergency medicine specialists independently evaluate each anonymized note, blinded to the original physician diagnosis and to each other. Primary diagnosis assigned by at least 2/3 specialists (majority vote) constitutes the gold standard. A 5-case calibration session precedes the main evaluation.

DATA PRIVACY: All anamnesis notes are fully de-identified (name, ID number, date of birth, physician name removed) prior to processing. De-identified notes are stored in a password-protected encrypted database. Only de-identified text is transmitted to LLM APIs - no personal health data. Compliant with Turkish Personal Data Protection Law (KVKK No. 6698).

PATIENT AND PUBLIC INVOLVEMENT: Not applicable. This retrospective study uses fully anonymized existing records; no patient or public involvement in design or conduct.

DATA SHARING: Anonymized dataset will be shared via Zenodo upon article acceptance. Statistical analysis code (Jamovi project files and Python prompt scripts) will be available on GitHub.

Первичные конечные точки

  • Diagnostic Accuracy of GPT-4o for ICD-10 Chapter-Level Diagnosis [Срок оценки: At the time of single-session algorithmic evaluation (each case evaluated once following data extraction in June 2026).]
  • Diagnostic Accuracy of Claude 4.6 Sonnet for ICD-10 Chapter-Level Diagnosis [Срок оценки: At the time of single-session algorithmic evaluation (each case evaluated once following data extraction in June 2026).]
Вторичные конечные точки (5)
  • Cohen's Kappa Between GPT-4o Primary Diagnosis and Gold Standard [Срок оценки: At the time of algorithmic evaluation (June-July 2026)]
  • Cohen's Kappa Between Claude 4.6 Sonnet Primary Diagnosis and Gold Standard [Срок оценки: At the time of algorithmic evaluation (June-July 2026)]
  • Top-3 Diagnostic Accuracy of GPT-4o [Срок оценки: At the time of algorithmic evaluation (June-July 2026)]
  • Top-3 Diagnostic Accuracy of Claude 4.6 Sonnet [Срок оценки: At the time of algorithmic evaluation (June-July 2026)]
  • Treating Physician Diagnostic Accuracy Against Gold Standard [Срок оценки: At the time of the original clinical encounter (retrospective data spanning August-December 2025)]

Критерии участия

Критерии включения

  • Adult patients (aged 18 years and older) presenting to the emergency department.
  • Complete electronic health record available in the hospital information system (HBYS) containing a detailed anamnesis note with chief complaint, symptom duration, associated symptoms, and relevant medical history.
  • A definitive primary diagnosis recorded by the treating emergency physician using ICD-10 codes at the time of patient file closure.

Критерии исключения

  • Emergency department anamnesis notes containing fewer than 50 words or completely lacking substantive clinical content\[cite: 1\].
  • Pediatric cases (age under 18 years)\[cite: 1\].
  • Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index \[ESI\] level 1)\[cite: 1\].
  • Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations\[cite: 1\].
  • Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry\[cite: 1\].

Критерии приведены из реестра в оригинале (на английском). Окончательную оценку соответствия проводит исследовательский центр.

Здоровые добровольцы: Нет

Дизайн исследования

Модель наблюдения
Когортное

Центры проведения

Turkey (Türkiye) · 1 центр
  • Marmara University Pendik Training and Research Hospital — Istanbul

Публикации

  • Newman-Toker DE, Peterson SM, Badihian S, Hassoon A, Nassery N, Parizadeh D, Wilson LM, Jia Y, Omron R, Tharmarajah S, Guerin L, Bastani PB, Fracica EA, Kotwal S, Robinson KA. Diagnostic Errors in the Emergency Department: A Systematic Review [Internet]. Rockville (MD): Agency for Healthcare Research and Quality (US); 2022 Dec. Report No.: 22(23)-EHC043. Available from http://www.ncbi.nlm.nih.gov/ PMID 36574484
  • Wei J et al. Chain-of-thought prompting elicits reasoning in LLMs. NeurIPS. 2022;35:24824-24837.
  • Sounderajah V, Guni A, Liu X, Collins GS, Karthikesalingam A, Markar SR, Golub RM, Denniston AK, Shetty S, Moher D, Bossuyt PM, Darzi A, Ashrafian H; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025 Oct;31(10):3283-3289. doi: 10.1038/s41591-025-03953-8. Epub 2025 Sep 15. PMID 40954311
  • Bossuyt PM, Reitsma JB, Bruns DE, Gatsonis CA, Glasziou PP, Irwig L, Lijmer JG, Moher D, Rennie D, de Vet HC, Kressel HY, Rifai N, Golub RM, Altman DG, Hooft L, Korevaar DA, Cohen JF; STARD Group. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015 Oct 28;351:h5527. doi: 10.1136/bmj.h5527. PMID 26511519
  • Niset A, Melot I, Pireau M, Englebert A, Scius N, Flament J, El Hadwe S, Al Barajraji M, Thonon H, Barrit S. Grounded large language models for diagnostic prediction in real-world emergency department settings. JAMIA Open. 2025 Oct 21;8(5):ooaf119. doi: 10.1093/jamiaopen/ooaf119. eCollection 2025 Oct. PMID 41127256
  • Williams CYK, Miao BY, Kornblith AE, Butte AJ. Evaluating the use of large language models to provide clinical recommendations in the Emergency Department. Nat Commun. 2024 Oct 8;15(1):8236. doi: 10.1038/s41467-024-52415-1. PMID 39379357
  • Hoppe JM, Auer MK, Struven A, Massberg S, Stremmel C. ChatGPT With GPT-4 Outperforms Emergency Department Physicians in Diagnostic Accuracy: Retrospective Analysis. J Med Internet Res. 2024 Jul 8;26:e56110. doi: 10.2196/56110. PMID 38976865
  • Kanjee Z, Crowe B, Rodman A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA. 2023 Jul 3;330(1):78-80. doi: 10.1001/jama.2023.8288. PMID 37318797

Идентификаторы

NCT: NCT07632859 · 09.2026.26-0514

Первоисточники (государственные реестры)

Открыть это исследование на ClinicalTrials.gov ↗