Diagnostic Superiority in Radiology: A Comparative Analysis of Gemini Pro and ChatGPT-4 for Chest X-Ray Interpretation
Keywords:
Artificial Intelligence; Chest X-Ray; ChatGPT-4; Gemini Pro; RadiologyAbstract
Background: Multimodal generative artificial intelligence is increasingly being explored as a potential support technology for radiological image interpretation. However, differences in diagnostic performance between commercially available AI systems remain insufficiently characterized.
Objective: This study compared the diagnostic performance of Gemini Pro and ChatGPT-4 in interpreting chest X-ray examinations.
Methods: A quantitative, comparative, cross-sectional diagnostic design was applied to 60 chest X-ray examinations. The same radiographs were independently assessed by Gemini Pro and ChatGPT-4 using a standardized prompt. Their interpretations were compared with an established reference diagnosis. Accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1-score, Cohen's kappa, and McNemar's test were used for evaluation.
Results: Gemini Pro achieved 78.3% overall normal–abnormal classification accuracy, 16.7% sensitivity, and 85.2% specificity. ChatGPT-4 achieved 53.3% accuracy, 66.7% sensitivity, and 51.9% specificity. For specific diagnostic identification, ChatGPT-4 achieved 58.3% accuracy compared with 18.3% for Gemini Pro. The paired comparison demonstrated a statistically significant difference between the models (McNemar's χ² = 18.893, p < .001).
Conclusion: The findings demonstrate distinct performance profiles between the two AI systems. ChatGPT-4 showed stronger sensitivity and specific diagnostic recognition, whereas Gemini Pro demonstrated better binary classification accuracy and specificity. Both systems require further validation using larger, balanced datasets before clinical implementation.
