Breast cancer AI test finds no single chatbot excels across every measure

A head-to-head comparison of five leading LLMs tested what happens when breast cancer questions move from textbook knowledge to clinical cases and the everyday concerns patients bring to their care teams.

A comparative study of large language models in responding to breast cancer–related questions. Image Credit: Lightspring / Shutterstock

A comparative study of large language models in responding to breast cancer–related questions. Image Credit: Lightspring / Shutterstock

In a recent study published as an 'Article in Press' in the journal Scientific Reports, researchers conducted a cross-sectional comparison to evaluate the performance of modern large language models (LLMs) in providing support for breast cancer health information.

The study tested ChatGPT-5.2, ChatGPT-4o, Gemini 3.0, DeepSeek, and ERNIE Bot using 90 multiple-choice questions, 10 de-identified clinical cases, and responses to 20 common patient concerns.

The study’s investigations revealed that while all five models showed high accuracy on standardized breast cancer knowledge questions, ranging from 82.22% to 94.44%, they differed in linguistic complexity and expert-rated quality measures.

Human experts gave ChatGPT-5.2 and DeepSeek the highest completeness scores, whereas Gemini 3.0 received the highest expert-rated readability score. DeepSeek also produced the lowest reading-difficulty score in the automated Chinese-language assessment.

The study concluded that no single model could be considered the ‘best’ in answering breast cancer questions and recommended evaluating LLMs across multiple measures, with future studies incorporating patient-based assessments and real clinical settings.

Background

Global estimates for 2020 indicate that ~685,000 women died from breast cancer, accounting for about 16% of female cancer deaths worldwide and highlighting the disease’s global public health burden.

While advances in breast cancer diagnosis and treatment have improved survival, postoperative physical changes and side effects of chemotherapy and radiation therapy can contribute to psychological distress, which may include fatigue, anxiety, and depression.

Breast cancer patients increasingly seek additional health information about their condition, treatment, recovery, and psychological concerns, with credible and understandable information helping them take a more active role in daily care and clinical decisions.

While modern conversational large language models can generate personalized health content, their accuracy, completeness, helpfulness, safety, readability, and linguistic complexity remain areas of active study in the context of breast cancer information support.

About the study

The present study aimed to provide a reference for evaluating LLM use in breast cancer health information by systematically comparing the five selected models. These models were assessed through their official web interfaces using default settings between December 15, 2025, and January 15, 2026. All prompts were entered in simplified Chinese.

The study evaluated LLM performance across three primary tasks. First, standardized knowledge was assessed using 90 multiple-choice questions from textbooks and clinical guidelines.

Next, a clinical case analysis was conducted in which LLMs were provided with 10 de-identified patient cases across five clinical domains (diagnosis, treatment, postoperative care, psychosocial support, and prognosis assessment and rehabilitation).

Finally, the study evaluated LLM responses to patient concerns, assessing 20 common clinical questions repeated across three sessions to measure output consistency and stability. The Chinese-language reading difficulty of LLM-generated text was objectively quantified using the Ludong University Text Grading Platform (LDU-TGP).

Three breast cancer specialists (‘human experts’) blindly and independently rated LLM responses for completeness, correctness, readability, helpfulness, and safety using a 5-point Likert scale. Statistical analyses compared model accuracy and expert ratings. Agreement among the three expert raters was poor overall, so individual ratings were retained separately in the analysis rather than averaged.

Study findings

The study’s standardized knowledge evaluations revealed that all models performed with a high degree of medical accuracy. DeepSeek achieved the highest numerical accuracy (94.44%; 85 of 90 correct), followed by ChatGPT-5.2 and ERNIE Bot at 88.89% (80 of 90 correct), Gemini 3.0 at 87.78% (79 of 90 correct), and ChatGPT-4o at 82.22% (74 of 90 correct).

Although the models differed overall in standardized knowledge performance, adjusted pairwise comparisons did not reveal a statistically significant difference between any pair of models. The findings, therefore, did not establish that the models were equivalent.

LDU-TGP evaluations demonstrated that LLMs varied substantially based on the reading complexity of their responses. ChatGPT-5.2 was found to produce the most complex prose (mean = 21.22; higher is worse) while DeepSeek generated the least linguistically difficult and most consistent Chinese-language case-analysis text according to this automated measure (mean = 13.15).

Expert-evaluated scores differed by model across all five dimensions: completeness, correctness, readability, helpfulness, and safety. ChatGPT-5.2 was found to achieve the highest completeness score (mean = 4.350 out of 5.0) while Gemini 3.0 (mean = 4.108) and DeepSeek (mean = 4.075) led in the readability and safety dimensions, respectively.

ERNIE Bot received the lowest estimated mean scores across all five expert-rated dimensions, including completeness (mean = 3.583 out of 5.0).

Conclusions

The study indicates that the five web-hosted LLMs tested performed similarly overall on standardized breast cancer knowledge, but differed in linguistic complexity and expert-rated response quality.

For example, while ChatGPT-5.2 and DeepSeek received higher expert-rated completeness scores, Gemini 3.0 received the highest expert-rated readability score. The study did not test patient comprehension or the safety and effectiveness of these models in direct patient-world clinical decision-making. All 10 clinical cases also came from a single hospital, which may limit the generalizability of the findings. The authors called for patient-based assessments and testing in actual clinical settings before the findings are applied to practice.

Journal reference:
Hugo Francisco de Souza

Written by

Hugo Francisco de Souza

Hugo Francisco de Souza is a scientific writer based in Bangalore, Karnataka, India. His academic passions lie in biogeography, evolutionary biology, and herpetology. He is currently pursuing his Ph.D. from the Centre for Ecological Sciences, Indian Institute of Science, where he studies the origins, dispersal, and speciation of wetland-associated snakes. Hugo has received, amongst others, the DST-INSPIRE fellowship for his doctoral research and the Gold Medal from Pondicherry University for academic excellence during his Masters. His research has been published in high-impact peer-reviewed journals, including PLOS Neglected Tropical Diseases and Systematic Biology. When not working or writing, Hugo can be found consuming copious amounts of anime and manga, composing and making music with his bass guitar, shredding trails on his MTB, playing video games (he prefers the term ‘gaming’), or tinkering with all things tech.

Citations

Please use one of the following formats to cite this article in your essay, paper or report:

  • APA

    Francisco de Souza, Hugo. (2026, September 10). Breast cancer AI test finds no single chatbot excels across every measure. News-Medical. Retrieved on September 10, 2026 from https://www.news-medical.net/news/20260910/Breast-cancer-AI-test-finds-no-single-chatbot-excels-across-every-measure.aspx.

  • MLA

    Francisco de Souza, Hugo. "Breast cancer AI test finds no single chatbot excels across every measure". News-Medical. 10 September 2026. <https://www.news-medical.net/news/20260910/Breast-cancer-AI-test-finds-no-single-chatbot-excels-across-every-measure.aspx>.

  • Chicago

    Francisco de Souza, Hugo. "Breast cancer AI test finds no single chatbot excels across every measure". News-Medical. https://www.news-medical.net/news/20260910/Breast-cancer-AI-test-finds-no-single-chatbot-excels-across-every-measure.aspx. (accessed September 10, 2026).

  • Harvard

    Francisco de Souza, Hugo. 2026. Breast cancer AI test finds no single chatbot excels across every measure. News-Medical, viewed 10 September 2026, https://www.news-medical.net/news/20260910/Breast-cancer-AI-test-finds-no-single-chatbot-excels-across-every-measure.aspx.

Comments

The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of News Medical.
Post a new comment
Post

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.

You might also like...
Cancer-fighting T cells need small amounts of ROS to attack tumors, study finds