---
title: "ChatGPT versus Gemini for Patient-Facing Breast Cancer Questions: Comparative Evaluation"
id: "plos-one-13-comparative-evaluation-of-chatgpt-and-gemini-responses-to-patient-oriented"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-13-comparative-evaluation-of-chatgpt-and-gemini-responses-to-patient-oriented"
content_type: "clinical_feed_article"
specialty: "Oncology"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358325"
published_at: "2026-09-16T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# ChatGPT versus Gemini for Patient-Facing Breast Cancer Questions: Comparative Evaluation
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-13-comparative-evaluation-of-chatgpt-and-gemini-responses-to-patient-oriented
- **Specialty:** [Oncology](https://medichelpline.com/clinical-feed/oncology.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358325)
- **Published At:** 2026-09-16T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- This study compared responses from **ChatGPT (GPT-5)** and **Gemini 2.5 Flash** to 40 frequently asked, patient-oriented questions about **breast cancer** and its surgical management. - Questions were collected from Turkish online sources (Google, YouTube, and institutional FAQ pages) and grouped into common categories such as symptoms, diagnosis, risk factors, imaging, prognosis, treatments, and side effects. - Both models were queried under identical conditions on August 13, 2025, in incognito mode; each question was asked in a new session and responses were anonymized for blinded assessment. - Four experienced general surgeons in breast surgery independently rated each anonymized response on three dimensions using a detailed five-point Likert scale: **scientific accuracy**, **clarity**, and **unnecessary detail**. - ChatGPT scored significantly higher than Gemini in scientific accuracy (median 4.75 vs. 4.25), clarity (median 4.75 vs. 4.25), and minimizing unnecessary detail (median 5.00 vs. 4.50); overall median scores were 4.83 for ChatGPT and 4.33 for Gemini. - Gemini produced significantly longer responses than ChatGPT (mean 244 ± 84 vs. 170 ± 47 words). - Statistical methods: Shapiro–Wilk for distribution; Wilcoxon signed-rank test for nonparametric paired comparisons; paired t-test for response length; effect sizes reported with Cohen’s d; Hodges–Lehmann for median differences; analyses done in SPSS v24. - Authors concluded both models had acceptable performance for patient education, but ChatGPT’s higher ratings and shorter, more focused replies suggest greater efficiency. They emphasized that neither model should be used for diagnosis or clinical decisions and that AI health information requires expert supervision and is best limited to educational/supportive roles. - The paper is open access in PLOS ONE (published September 16, 2026); all underlying data and supplementary tables were reported as available within the manuscript and supporting information.
## Clinical Analysis & Structured Key Points
Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Reader Comments Figures Figures Abstract Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-oriented questions about breast cancer and its surgical treatment regarding scientific accuracy, clarity, and unnecessary detail. Methods Forty frequently asked questions were collected from Turkish online sources. Both models were queried under identical conditions on August 13, 2025, and the responses were anonymized for blinded evaluation. Four general surgeons experienced in breast surgery independently assessed each response using a five-point Likert scale across three domains: scientific accuracy, clarity, and unnecessary detail. Results Response lengths were recorded and compared. ChatGPT achieved significantly higher median scores than Gemini in scientific accuracy [4.75 (3.50–5.00) vs. 4.25 (3.50–5.00); p < 0.001], clarity [4.75 (3.50–5.00) vs. 4.25 (3.25–5.00); p = 0.005], and unnecessary detail [5.00 (5.00–5.00) vs. 4.50 (3.50–5.00); p < 0.001]. The overall median score was 4.83 (4.08–5.00) for ChatGPT and 4.33 (3.50–4.92) for Gemini (p < 0.001). Gemini’s responses were significantly longer (244 ± 84 vs. 170 ± 47 words; p < 0.001). Conclusion Both models demonstrated acceptable overall performance. However, ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes. Citation: Yilmaz Bozok Y, Acar N, Atahan MK, Karaali C (2026) Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer. PLoS One 21(9): e0358325. https://doi.org/10.1371/journal.pone.0358325 Editor: Lorenzo Faggioni, University of Pisa, ITALY Received: January 14, 2026; Accepted: August 31, 2026; Published: September 16, 2026 Copyright: © 2026 Yilmaz Bozok et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: All data underlying the findings of this study are fully available within the manuscript and its Supporting Information file (Supplementary Tables 1–8). Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. Introduction Breast cancer, the most common malignancy among women, remains a major public health concern in both developed and developing countries [ 1 ]. Despite significant technological advances in screening and treatment, it continues to be one of the leading causes of cancer-related mortality among women worldwide [ 2 ]. Globally, there is a marked inequality in the burden of breast cancer. Although the incidence is higher in high-income countries, the mortality rate is lower due to the effectiveness of early detection programs and better access to treatment. In contrast, low- and middle-income countries experience increasing mortality rates, largely attributable to diagnostic delays and limited healthcare resources [ 3 ]. Delays in the diagnosis and treatment of breast cancer have been associated with poor quality of life and reduced survival outcomes [ 4 , 5 ]. A multicenter study conducted in Türkiye focusing on the causes of diagnostic and treatment delays in breast cancer identified key patient-related factors, including lack of knowledge (12.4%), fear of losing the breast (8.9%), and fear of death (9.8%). These findings highlight that patient-related factors play a predominant role in diagnostic delays [ 6 ]. Access to accurate medical information is crucial for reducing patient anxiety and guiding individuals toward timely medical consultation. In this context, large language models (LLMs) have recently emerged as modern tools that can potentially enhance health literacy and support patient education. Health literacy plays a critical role in the effectiveness of patient education materials. Previous recommendations from the National Institutes of Health, the US Department of Health and Human Services, and the American Medical Association suggest that patient-directed educational materials should ideally be written at approximately a sixth-grade reading level to improve accessibility and comprehension [ 7 ]. Previous reports suggest that a substantial proportion of internet users seek medical information online, with rates approaching 80% in some studies [ 8 , 9 ]. In addition, ChatGPT was estimated to have nearly 300 million weekly active users and approximately 3.8 billion monthly visits as of November 2024 [ 10 ]. AI-generated medical information may increasingly influence patient understanding, treatment expectations, and healthcare-related decision-making processes [ 11 , 12 ]. However, their rapid proliferation has raised important questions regarding the accuracy and reliability of the information they provide. Growing concerns exist that AI-based systems may generate or disseminate misleading health information at a large scale [ 13 ]. In recent years, bibliometric and Altmetric analyses have increasingly been used to evaluate the scientific visibility and online dissemination of medical information. These approaches provide complementary insights into the relationship between academic publications, social media engagement, and public interest in health-related topics [ 14 ]. The aim of this study was to comparatively evaluate the responses generated by ChatGPT and Gemini to patient-oriented questions on breast cancer and its surgical management in terms of scientific accuracy, clarity, and the level of unnecessary detail. Methods This cross-sectional comparative study evaluated patient-oriented questions identified through Google and YouTube searches in the local language and by reviewing the “frequently asked questions” sections of health institutions. When compiling the question list related to breast cancer and its surgery, the following categories were considered: symptoms and diagnosis, risk factors and prevention, imaging methods, disease course and prognosis, treatment options, and treatment-related side effects. After merging and simplifying similar questions, a final list consisting of 40 questions was obtained. In this study, two LLMs, ChatGPT (GPT 5, OpenAI, USA) and Gemini 2.5 Flash (Google DeepMind, United Kingdom), were used. To standardize the evaluation conditions, separate user accounts were created for each LLM, and all responses were obtained within the same date and time range (August 13, 2025, 17:00–20:00). In addition, searches and model interactions were performed using incognito mode to minimize potential personalization effects related to previous browsing activity or search history. The questions were directed to each model without any additional prompts or role instructions, and each question was asked in a new session. This approach prevented any possible contextual influence from preceding or subsequent questions. The responses were transferred to anonymized booklets by the study designer without any modification, ensuring blinded evaluation by the raters. The second booklet was delivered to the raters at least seven days after the first one was completed to minimize potential recall bias. In addition, the word count of each response generated by the LLMs was recorded for every question. Four general surgery specialists experienced in breast cancer and breast surgery, including two professors and two associate professors, participated as raters. The AI-generated responses were assessed across three main dimensions: scientific accuracy, clarity, and level of unnecessary detail. A five-point Likert-type scale (1–5) was used for each dimension. The scale items were clearly defined to describe the meaning of each level in detail. For example, a score of “1” represented responses containing major scientific errors, unclear or confusing language, or excessive irrelevant information; whereas a score of “5” indicated responses fully consistent with current guidelines, clear, coherent, and limited to essential information. These detailed definitions were included in the evaluation form to improve consistency among raters. As this research focused solely on the analysis of AI-generated content and did not involve any patient data, medical records, archived samples, or identifiable personal information, ethics committee approval and informed consent were not required. Statistical analysis The distribution characteristics of the data were evaluated using the Shapiro–Wilk test. Since the score-related variables did not show normal distribution, paired comparisons were performed using the non-parametric Wilcoxon signed-rank test. These variables were expressed as median (minimum–maximum) values. When a significant difference was found and the median values were equal, the direction of the difference was interpreted based on the comparison of mean ranks. The Wilcoxon signed-rank test was used for paired groups; effect size was reported using Cohen's d coefficient. Median differences between paired measurements and their 95% confidence intervals were estimated using the Hodges–Lehmann method. The response length data showed normal distribution; therefore, comparisons were performed using the paired-samples t-test. The effect size was reported using Cohen's d coefficient. All analyses were performed using IBM SPSS Statistics version 24.0 (IBM Corp., Armonk, NY, USA). A p value of <0.05 was considered statistically significant. Results Based on the aggregated scores of four raters, the median score of ChatGPT for scientific accuracy was significantly higher than that of Gemini [4.75 (3.50–5.00) vs. 4.25 (3.50–5.00), p < 0.001]. In the clarity dimension, the median score of ChatGPT [4.75 (3.50–5.00)] was also higher than that of Gemini [4.25 (3.25–5.00), p = 0.005]. Similarly, in the level of unnecessary detail, ChatGPT showed a significantly higher median score [5.00 (5.00–5.00) vs. 4.50 (3.50–5.00), p < 0.001]. Comparisons of scientific accuracy, clarity, and level of unnecessary detail are shown in Fig 1 . Download: PNG larger image TIFF original image Fig 1. Comparison of ChatGPT and Gemini across three evaluation domains: scientific accuracy (Acc), clarity (Clar), and unnecessary detail (Det), as rated by four evaluators (R1–R4). Bars show median scores with error bars for variability. https://doi.org/10.1371/journal.pone.0358325.g001 When overall mean scores were analyzed, the median score of ChatGPT was 4.83 (4.08–5.00), while that of Gemini was 4.33 (3.50–4.92), and the difference was statistically significant (p < 0.001). Detailed comparisons for each evaluation domain and rater are presented in Table 1 . Gemini generated significantly longer responses than ChatGPT (244 ± 84 vs. 170 ± 47 words, p < 0.001). Detailed comparisons of response lengths are presented in Table 2 . Download: PNG larger image TIFF original image Table 1. Comparison of ChatGPT and Gemini across three evaluation domains: scientific accuracy, clarity, and level of unnecessary detail, for each evaluator (R1–R4) and overall results. https://doi.org/10.1371/journal.pone.0358325.t001 Download: PNG larger image TIFF original image Table 2. Word count comparison of ChatGPT and Gemini responses for 40 breast cancer questions. https://doi.org/10.1371/journal.pone.0358325.t002 For Rater 1, no statistically significant difference was observed between ChatGPT and Gemini in scientific accuracy (p = 0.519) or clarity (p = 0.835). However, ChatGPT scored significantly higher in the level of unnecessary detail [median 5.00 (5.00–5.00) vs. 4.00 (3.00–5.00), p < 0.001]. The overall median scores of Rater 1 were 4.33 (3.67–5.00) for ChatGPT and 4.00 (3.33–5.00) for Gemini, with a statistically significant difference (p = 0.014). For Rater 2, ChatGPT achieved higher median scores than Gemini across all dimensions. In scientific accuracy [5.00 (4.00–5.00) vs. 4.00 (3.00–5.00)], clarity [5.00 (4.00–5.00) vs. 4.00 (3.00–5.00)], and level of unnecessary detail [5.00 (5.00–5.00) vs. 4.00 (3.00–5.00)] the differences were statistically significant (p < 0.001, p < 0.001, p = 0.003 respectively). The overall median scores of Rater 2 were 5.00 (4.33–5.00) for ChatGPT and 4.00 (3.00–5.00) for Gemini (p < 0.001). For Rater 3, a significant difference favoring ChatGPT was found in scientific accuracy (p = 0.003). Although the medians were equal [5.00 (4.00–5.00) vs. 5.00 (4.00–5.00)], the difference was supported by mean rank analysis. No significant difference was found in clarity [5.00 (4.00–5.00) vs. 5.00 (3.00–5.00); p = 0.251]. In the level of unnecessary detail, as in scientific accuracy, equal medians [5.00 (5.00–5.00) vs. 5.00 (4.00–5.00)] but higher mean ranks indicated a difference favoring ChatGPT (p = 0.025). The overall median scores of Rater 3 were 5.00 (4.33–5.00) for ChatGPT and 4.66 (3.67–5.00) for Gemini (p = 0.011). For Rater 4, ChatGPT demonstrated higher performance than Gemini in scientific accuracy [5.00 (3.00–5.00) vs. 4.00 (3.00–5.00); p = 0.032]. No statistically significant difference was observed in clarity, as the median values were equal for both models [5.00 (3.00–5.00) vs. 5.00 (3.00–5.00); p = 0.375]. In terms of the level of unnecessary detail, although the median scores were equal [5.00 (5.00–5.00) vs. 5.00 (4.00–5.00)], mean rank analysis indicated a statistically significant difference favoring ChatGPT (p = 0.000532). Overall, the median score was higher for ChatGPT than for Gemini [5.00 (3.67–5.00) vs. 4.66 (3.33–5.00); p = 0.008]. Discussion This study comprehensively evaluated artificial intelligence-generated responses to patient-oriented questions regarding breast cancer and its surgical management. Both models demonstrated overall high performance; however, ChatGPT outperformed Gemini significantly in terms of scientific accuracy, clarity, level of unnecessary detail, and overall mean score. As in many other fields, artificial intelligence applications are rapidly expanding within healthcare [ 15 ]. AI assistance now extends from diagnostic and imaging processes to reducing clinicians’ workloads and enhancing patient engagement. Despite these advantages and the trust these systems often inspire, the possibility of errors and the uncontrolled spread of misinformation remain among the most critical concerns in medical contexts. Tan et al. evaluated ChatGPT’s responses to patient questions about glaucoma and emphasized that, despite promising results, the potential for generating incorrect or misleading information may pose risks for patients [ 16 ]. Similarly, Yalamanchili et al. found that while large language models could serve as an effective alternative to online health resources for radiation oncology patients, their readability remained suboptimal [ 17 ]. Similarly, a recent comparative study evaluating ChatGPT, Gemini, and Perplexity reported concerns regarding the reliability of AI-generated medical information and emphasized that these systems may still produce factually inaccurate or difficult-to-read responses [ 18 ]. In another study focusing on lung cancer and surgical management, ChatGPT achieved scores above 4.5 across all dimensions of scientific adequacy, clarity, and accuracy [ 19 ]. However, the authors cautioned that the model’s reliance on outdated training data and the potential for “hallucinations” (fabricated or incorrect information) may result in misinformation. In contrast, our study identified only minor phrasing differences or trivial factual inaccuracies, with no clinically significant misinformation detected. Nevertheless, the possibility of such errors, particularly in the context of medical decision-making, should not be underestimated given their potential clinical implications. Another key finding was that Gemini produced significantly longer responses compared to ChatGPT, which may represent a potential disadvantage. For instance, for the question “Can men get breast cancer?”, ChatGPT generated a 138-word response, whereas Gemini produced 514 words. While Gemini’s overall performance remained acceptable, the longer and less focused responses may have contributed to its lower scores. Similar findings have been reported in previous Turkish-language comparative studies using the same LLMs in different medical fields, where Gemini consistently generated longer responses than GPT-based models [ 20 , 21 ]. Tong et al. reported opposite findings in a different domain, noting that Gemini generated shorter responses than ChatGPT [ 22 ]. These inconsistencies suggest that response length variations depend on factors such as model version, subject domain, and evaluation conditions. The observed performance differences may also be related to differences in the models’ development processes, technological maturity, and response-generation characteristics. However, given that this comparison was conducted within a specific language, clinical domain, and expert evaluator group, the observed superiority should not be assumed to persist under different evaluation conditions. To improve readability and prevent information overload, integrating response-length control mechanisms, such as word or character limits, into user interfaces could help optimize the outputs of these systems. Our study contributes to the limited literature evaluating AI-generated responses to patient questions about breast cancer and its surgical treatment. Piao et al. conducted a similar study in China comparing ChatGPT, ERNIE Bot, and ChatGLM and found comparable overall performances among the models, although ChatGPT was rated higher for human-like empathy [ 23 ]. The authors emphasized that such tools are suitable for addressing general informational queries but remain unreliable for complex clinical topics. Likewise, Roldan-Vasquez et al. found ChatGPT’s responses to breast surgery–related questions to be accurate and reliable, with only minor, clinically insignificant errors [ 24 ]. To the best of our knowledge, this study represents one of the first comparative evaluations of large language model responses to patient-oriented questions regarding breast cancer and its surgical management in the Turkish language. In addition, the blinded evaluation performed by multiple specialists under standardized assessment conditions represents a methodological strength of the study. This study has several limitations. First, patient questions were collected exclusively in the local language (Turkish), which may limit cross-linguistic comparability. Previous studies evaluating Turkish internet-based patient education materials have reported important limitations regarding the readability, quality, and reliability of online health information [ 25 ]. Therefore, the findings of the present Turkish-language study may not be directly generalizable to other languages and healthcare settings. Future studies conducted in multiple languages with similar designs could enhance the generalizability of findings. Second, large language models are continuously updated and may produc
## Related Clinical Research

- [GSK Highlights Lung Cancer Data From Two Drugs — STAT+ Briefing](https://medichelpline.com/clinical-feed/stat-news-0-stat-gsk-touts-lung-cancer-data-from-two-drugs.md)
- [Daraxonrasib shows activity in RAS‑mutant NSCLC after FDA approval in pancreatic cancer](https://medichelpline.com/clinical-feed/medical-news-today-0-pancreatic-cancer-drug-may-also-help-fight-treatment-resistant-lung-cancer.md)
- [Sen. Warren demands access to Trump drug-pricing contracts, accuses RFK Jr. of concealing them](https://medichelpline.com/clinical-feed/stat-news-0-stat-sen-warren-demands-to-see-trump-drug-pricing-pharma-contracts-accuses-rfk.md)
- [Carbon monoxide exposure may explain lower Parkinson's risk linked to smoking](https://medichelpline.com/clinical-feed/medical-news-today-0-why-smokers-may-have-a-lower-parkinson-s-risk-it-s-not-the-nicotine.md)
- [Self-acupressure for fatigue-sleep disturbance-depression cluster in breast cancer survivors: phas](https://medichelpline.com/clinical-feed/pubmed-42749858.md) (DOI: 10.1007/s00520-026-11155-2)

## Navigation
- [← Back to Oncology Feed](https://medichelpline.com/clinical-feed/oncology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.