15. International Trakya Family Medicine Congress

16-19 April 2026, Balkan Congress Center, Edirne

Diagnostic Performance of ChatGPT in Shoulder Dislocation Detection on Pre-Reduction Radiographs

Çağdaş Derdiyok, Mahmut Sami Oflaz

Keywords: Shoulder dislocation, artificial intelligence, large language model, ChatGPT, diagnostic imaging, emergency department

Aim:

This study aimed to evaluate the diagnostic performance of ChatGPT in detecting shoulder dislocation on pre-reduction shoulder radiographs in the emergency department, using orthopedic specialist evaluation as the reference standard.

Method:

In this retrospective diagnostic accuracy study, data of 179 patients (aged 18+) who presented to the Emergency Department with suspected shoulder dislocation between 2023-2025 were reviewed. Shoulder radiographs were evaluated independently and blindly by an orthopedic specialist and ChatGPT (GPT-4V). Primary endpoints were sensitivity and specificity; secondary endpoints were PPV, NPV, accuracy, and kappa agreement.

Results:

Shoulder dislocation was detected in 133 patients (74.3%) by the orthopedic specialist (121 anterior, 12 posterior). ChatGPT showed sensitivity of 93.2% (95% CI: 89.0-97.5%), specificity of 97.8% (95% CI: 93.6-100.0%), PPV of 99.2%, NPV of 83.3%, and accuracy of 94.4% (95% CI: 91.0-97.8%). The kappa agreement was 0.862 (excellent). For associated fracture detection, sensitivity was 75.0% and specificity was 100.0%.

Conclusions:

ChatGPT demonstrates high sensitivity and specificity in diagnosing shoulder dislocation on pre-reduction shoulder radiographs. Large language models can be used as decision support tools for emergency physicians.

#45