This cross-sectional comparative study evaluated how two large language models, ChatGPT (GPT-5) and Gemini 2.5 Flash, respond to patient-oriented questions about breast cancer and its surgical management. The primary objective was to compare model outputs on three clinically relevant dimensions: scientific accuracy, clarity, and the amount of unnecessary detail provided. The authors framed the work against the need for accessible, reliable information to reduce diagnostic delay and support patient education.
Forty frequently asked questions were compiled from Turkish-language online sources, including Google and YouTube searches and frequently asked questions pages of health institutions. Questions covered categories such as symptoms and diagnosis, risk factors and prevention, imaging methods, disease course and prognosis, treatment options, and treatment-related side effects. Similar items were merged and simplified to create the final list of 40 questions used for model queries.
Responses were generated from ChatGPT (GPT-5, OpenAI) and Gemini 2.5 Flash (Google DeepMind). Separate user accounts were used and interactions occurred on August 13, 2025, between 17:00 and 20:00 to standardize timing. All sessions were performed in incognito mode to reduce personalization from prior browsing. Each question was submitted in a new session without additional role prompts or contextual instructions to avoid cross-question influence. Generated responses were collected without modification and anonymized prior to evaluation.
Four general surgery specialists experienced in breast cancer and breast surgery (two professors and two associate professors) acted as blinded raters. They independently evaluated each anonymized response across three dimensions: scientific accuracy, clarity, and level of unnecessary detail, using a detailed five-point Likert-type scale (1–5). Detailed definitions for each level were provided in the evaluation form. The second anonymized booklet of responses was delivered at least seven days after the first to reduce recall bias.
Data distribution was assessed using the Shapiro–Wilk test. Score-related variables were non-normally distributed and compared using the Wilcoxon signed-rank test for paired data; results were reported as medians with ranges. Median differences and 95% confidence intervals were estimated with the Hodges–Lehmann method. Response length data were normally distributed and compared with paired-samples t-tests. Effect sizes were reported using Cohen’s d. Analyses were conducted in IBM SPSS Statistics version 24. A p value < 0.05 was considered statistically significant.
Aggregated across four raters, ChatGPT achieved a significantly higher median score than Gemini for scientific accuracy: 4.75 (range 3.50–5.00) versus 4.25 (3.50–5.00), p < 0.001. In the clarity domain ChatGPT again scored higher: median 4.75 (3.50–5.00) versus 4.25 (3.25–5.00), p = 0.005. For unnecessary detail, ChatGPT’s median was 5.00 (5.00–5.00) compared with Gemini’s 4.50 (3.50–5.00), p < 0.001. The overall median score was 4.83 (4.08–5.00) for ChatGPT and 4.33 (3.50–4.92) for Gemini (p < 0.001).
These comparisons were visualized in the manuscript (Fig 1) and detailed per-rater and per-domain results appear in the supporting tables provided with the article.
Gemini produced significantly longer responses than ChatGPT. Mean response lengths reported were 244 ± 84 words for Gemini versus 170 ± 47 words for ChatGPT (statistical significance reported in the manuscript). The authors highlight that ChatGPT’s shorter, more focused replies corresponded with higher ratings for clarity and lower scores for unnecessary detail.
Readability metrics and recommended reading grade levels were discussed in the Introduction as contextual considerations for patient-directed materials; however, specific readability scores for the AI responses were not reported in the source.
The authors interpret the findings to mean that while both LLMs delivered acceptable performance for patient-oriented breast cancer questions, ChatGPT produced more accurate, clearer, and more concise answers in this dataset. They note the potential of LLMs to support health literacy and guide patients toward timely medical consultation, particularly given widespread online health information seeking. Nonetheless, they emphasize that AI-generated content may still propagate misleading information and therefore should not substitute for clinical evaluation or decision-making.
The study underscores that AI tools may be best used under expert supervision and deployed for educational or supportive purposes rather than diagnostic use. The authors reference prior guideline recommendations that patient educational materials be written at accessible reading levels, noting the broader relevance of clarity and conciseness when designing patient-facing AI outputs.
The study used a predefined set of 40 Turkish-language questions and evaluated two specific model versions at a single timepoint (August 13, 2025). Interactions were constrained to neutral prompts and single-session queries; other prompting strategies or model updates could affect performance. Rater assessment, while blinded and standardized, remains subjective even with detailed scoring rubrics. The manuscript reports that ethics approval and informed consent were not required because no patient data were used. The authors also stated that all underlying data and supplementary tables are available within the manuscript and supporting information.
Both evaluated LLMs demonstrated acceptable capability to answer common patient questions about breast cancer, but ChatGPT scored higher for scientific accuracy, clarity, and avoidance of unnecessary detail, while producing shorter responses. The authors conclude that these models, in their current forms, should not be used for diagnostic or clinical decision-making and that AI-generated health information should be applied under expert oversight for educational and supportive roles only.
Publication details: the study is published in PLOS ONE (received January 14, 2026; accepted August 31, 2026; published September 16, 2026). The authors reported no specific funding and no competing interests, and they stated that all data underlying the findings are available in the manuscript and supporting information.