Diagnostic Performance of ChatGPT in Shoulder Dislocation Detection on Pre-Reduction Radiographs
Çağdaş Derdiyok, Mahmut Sami Oflaz
Keywords: Shoulder dislocation, artificial intelligence, large language model, ChatGPT, diagnostic imaging, emergency department
Aim:
This study aimed to evaluate the diagnostic performance of ChatGPT in detecting shoulder dislocation on pre-reduction shoulder radiographs in the emergency department, using orthopedic specialist evaluation as the reference standard.
Method:
In this retrospective diagnostic accuracy study, data of 179 patients (aged 18+) who presented to the Emergency Department with suspected shoulder dislocation between 2023-2025 were reviewed. Shoulder radiographs were evaluated independently and blindly by an orthopedic specialist and ChatGPT (GPT-4V). Primary endpoints were sensitivity and specificity; secondary endpoints were PPV, NPV, accuracy, and kappa agreement.
Results:
Shoulder dislocation was detected in 133 patients (74.3%) by the orthopedic specialist (121 anterior, 12 posterior). ChatGPT showed sensitivity of 93.2% (95% CI: 89.0-97.5%), specificity of 97.8% (95% CI: 93.6-100.0%), PPV of 99.2%, NPV of 83.3%, and accuracy of 94.4% (95% CI: 91.0-97.8%). The kappa agreement was 0.862 (excellent). For associated fracture detection, sensitivity was 75.0% and specificity was 100.0%.
Conclusions:
ChatGPT demonstrates high sensitivity and specificity in diagnosing shoulder dislocation on pre-reduction shoulder radiographs. Large language models can be used as decision support tools for emergency physicians.
#45