Breast cancer AI test finds no single chatbot excels across every measure

by · News-Medical

A head-to-head comparison of five leading LLMs tested what happens when breast cancer questions move from textbook knowledge to clinical cases and the everyday concerns patients bring to their care teams.

A comparative study of large language models in responding to breast cancer–related questions. Image Credit: Lightspring / Shutterstock

Want to read later? Download your clean, ad-free, print-friendly PDF copy by clicking here.

The study tested ChatGPT-5.2, ChatGPT-4o, Gemini 3.0, DeepSeek, and ERNIE Bot using 90 multiple-choice questions, 10 de-identified clinical cases, and responses to 20 common patient concerns.

The study’s investigations revealed that while all five models showed high accuracy on standardized breast cancer knowledge questions, ranging from 82.22% to 94.44%, they differed in linguistic complexity and expert-rated quality measures.

Human experts gave ChatGPT-5.2 and DeepSeek the highest completeness scores, whereas Gemini 3.0 received the highest expert-rated readability score. DeepSeek also produced the lowest reading-difficulty score in the automated Chinese-language assessment.

The study concluded that no single model could be considered the ‘best’ in answering breast cancer questions and recommended evaluating LLMs across multiple measures, with future studies incorporating patient-based assessments and real clinical settings.

Background

Global estimates for 2020 indicate that ~685,000 women died from breast cancer, accounting for about 16% of female cancer deaths worldwide and highlighting the disease’s global public health burden.

While advances in breast cancer diagnosis and treatment have improved survival, postoperative physical changes and side effects of chemotherapy and radiation therapy can contribute to psychological distress, which may include fatigue, anxiety, and depression.

Breast cancer patients increasingly seek additional health information about their condition, treatment, recovery, and psychological concerns, with credible and understandable information helping them take a more active role in daily care and clinical decisions.

While modern conversational large language models can generate personalized health content, their accuracy, completeness, helpfulness, safety, readability, and linguistic complexity remain areas of active study in the context of breast cancer information support.

About the study

The present study aimed to provide a reference for evaluating LLM use in breast cancer health information by systematically comparing the five selected models. These models were assessed through their official web interfaces using default settings between December 15, 2025, and January 15, 2026. All prompts were entered in simplified Chinese.

The study evaluated LLM performance across three primary tasks. First, standardized knowledge was assessed using 90 multiple-choice questions from textbooks and clinical guidelines.

Next, a clinical case analysis was conducted in which LLMs were provided with 10 de-identified patient cases across five clinical domains (diagnosis, treatment, postoperative care, psychosocial support, and prognosis assessment and rehabilitation).

Finally, the study evaluated LLM responses to patient concerns, assessing 20 common clinical questions repeated across three sessions to measure output consistency and stability. The Chinese-language reading difficulty of LLM-generated text was objectively quantified using the Ludong University Text Grading Platform (LDU-TGP).

Three breast cancer specialists (‘human experts’) blindly and independently rated LLM responses for completeness, correctness, readability, helpfulness, and safety using a 5-point Likert scale. Statistical analyses compared model accuracy and expert ratings. Agreement among the three expert raters was poor overall, so individual ratings were retained separately in the analysis rather than averaged.

Study findings

The study’s standardized knowledge evaluations revealed that all models performed with a high degree of medical accuracy. DeepSeek achieved the highest numerical accuracy (94.44%; 85 of 90 correct), followed by ChatGPT-5.2 and ERNIE Bot at 88.89% (80 of 90 correct), Gemini 3.0 at 87.78% (79 of 90 correct), and ChatGPT-4o at 82.22% (74 of 90 correct).

Although the models differed overall in standardized knowledge performance, adjusted pairwise comparisons did not reveal a statistically significant difference between any pair of models. The findings, therefore, did not establish that the models were equivalent.

LDU-TGP evaluations demonstrated that LLMs varied substantially based on the reading complexity of their responses. ChatGPT-5.2 was found to produce the most complex prose (mean = 21.22; higher is worse) while DeepSeek generated the least linguistically difficult and most consistent Chinese-language case-analysis text according to this automated measure (mean = 13.15).

Expert-evaluated scores differed by model across all five dimensions: completeness, correctness, readability, helpfulness, and safety. ChatGPT-5.2 was found to achieve the highest completeness score (mean = 4.350 out of 5.0) while Gemini 3.0 (mean = 4.108) and DeepSeek (mean = 4.075) led in the readability and safety dimensions, respectively.

ERNIE Bot received the lowest estimated mean scores across all five expert-rated dimensions, including completeness (mean = 3.583 out of 5.0).

Conclusions

The study indicates that the five web-hosted LLMs tested performed similarly overall on standardized breast cancer knowledge, but differed in linguistic complexity and expert-rated response quality.

For example, while ChatGPT-5.2 and DeepSeek received higher expert-rated completeness scores, Gemini 3.0 received the highest expert-rated readability score. The study did not test patient comprehension or the safety and effectiveness of these models in direct patient-world clinical decision-making. All 10 clinical cases also came from a single hospital, which may limit the generalizability of the findings. The authors called for patient-based assessments and testing in actual clinical settings before the findings are applied to practice.

Journal reference: