Rapid advances in artificial intelligence have prompted investigation into the role of large language models as educational and clinical decision-support tools. This study directly compared the performance of two LLMs—ChatGPT-4 and Gemini—with chest-disease resident physicians in analyzing clinical scenarios, with the stated aim of evaluating their potential roles in medical education and clinical decision support.
The research was a cross-sectional, comparative study conducted at a tertiary-care university hospital. The participant group comprised 28 resident physicians working in the department of chest diseases. The two AI models (ChatGPT-4 and Gemini) were presented with the same clinical scenarios as the residents. Responses from both humans and AI were then evaluated by blinded experts.
Four clinical scenarios were selected to represent common and high-stakes conditions in chest disease practice: massive pulmonary embolism, chronic obstructive pulmonary disease (COPD), asthma, and severe pneumonia/sepsis. Scoring and evaluation referenced current guideline frameworks, explicitly including the Global Initiative for Chronic Obstructive Lung Disease (GOLD), the Global Initiative for Asthma (GINA), and standards from the American Thoracic Society.
Responses were scored by blinded experts against guideline-based expectations. The assessment emphasized structured knowledge elements such as disease classification, listing of contraindications, and differential diagnoses, as well as practical management tasks relevant to emergency settings. The blinded review process aimed to reduce bias in comparing AI outputs with resident responses.
Overall, the two AI models achieved significantly higher scores than the resident physicians on structured questions that tested theoretical knowledge, classification ability, and the ability to list contraindications (reported significance: P < 0.05). The AI models also offered a broader set of possibilities in differential diagnosis compared with residents.
Although AI outperformed residents on theoretically oriented and structured items, residents performed at levels similar to the AI models in tasks requiring immediate, practical action—specifically emergency interventions such as shock management. The authors noted a pattern in answer style: AI tended to produce broader, more comprehensive differential diagnoses and guideline-driven explanations, whereas resident responses were described as more "telegraphic" and practice-oriented, focusing on actionable steps.
The study authors interpret the findings to indicate substantial potential for ChatGPT and Gemini to serve as clinical decision-support systems and as educational assistants that can accelerate clinicians' access to theoretical knowledge and guideline-based recommendations. However, the authors caution against viewing these models as replacements for human clinical reasoning or for the hands-on decision-making required in emergency management. Instead, the recommendation is to position AI models as complementary tools that augment resident and physician performance.
The PubMed abstract reports the study design, participant number (28 residents), the four clinical scenarios used, the guideline references, and that blinded experts scored responses. The abstract does not provide detailed scoring rubrics, inter-rater reliability metrics, the precise magnitude of score differences, or granular breakdowns by scenario beyond the summary findings. Those full methodological and numerical details are not reported in the abstract and would require consultation of the full text available from the publisher (Galenos) for comprehensive appraisal.
In this controlled comparison, AI large language models (ChatGPT-4 and Gemini) outperformed chest-disease residents on structured, guideline-based knowledge tasks, while residents matched AI in emergency, results-focused management. The authors conclude that these AI models have a meaningful role as decision-support and educational aids but should complement rather than replace human clinicians, particularly for tasks requiring rapid, pragmatic clinical action. No conflicts of interest were declared. The study is indexed at PMID 41979097 and DOI 10.4274/ThoracResPract.2026.2026-1-2; the full text is available through Galenos Publishing for readers seeking the full dataset and methods.