Source study found
Story checked
Breast cancer AI test finds no single chatbot excels across every measure (opens in a new tab)
news-medical.net · 2026-09-10
Short answer
MixedMixed.
The claims we could check match the study, but some claims were not covered by the evidence reviewed.
- 2 supported
- 4 not covered
Checked against the study summary. The full text wasn't available, so some details couldn't be settled either way.
Share this check
The story
Breast cancer AI test finds no single chatbot excels across every measure
news-medical.net · 2026-09-10
The story’s checkable claims.
Read the original story (opens in a new tab)NewsLink checks it
Mixed
Every claim we could check holds up. Two of six claims match the study. This overall rating is based only on the claims we could check. Four claims the study doesn't address.
- 2 supported
- 4 not covered
The source study
A comparative study of large language models in responding to breast cancer–related questions
Evidence layer
Claim by claim
Each claim gets a verdict. Expand it to see the evidence directly below.
Reading mode
Scan verdicts. Open evidence only when needed.
Browse by verdict
6 claims in this storyShowing all 6 claimsChoose a verdict to focus the list.
Claim 1 of 6Not coveredThe study concluded that no single model could be considered the best for answering breast cancer questions and recommended evaluating LLMs across multiple measures, with future studies using patient-based assessments and real clinical settings.View evidenceHide evidence
Why this verdict
The conclusion that no single model was best across all assessed metrics and that LLMs should be evaluated across multiple dimensions is supported by the abstract profile. However, the added recommendation that future studies use patient-based assessments and real clinical settings is not present in the abstract-level profile, so that part is not verifiable at this evidence depth.
Study evidence
Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
Study evidence
Automated LDU-TGP analysis showed that LLM-generated breast cancer texts had varying linguistic complexities across models.
“Text readability was analyzed using LDU-TGP.”
Claim 2 of 6Not coveredThe study tested ChatGPT-5.2, ChatGPT-4o, Gemini 3.0, DeepSeek, and ERNIE Bot using 90 multiple-choice questions, 10 de-identified clinical cases, and 20 common patient concerns.View evidenceHide evidence
As stated90 multiple-choice questions; 10 clinical cases; 20 patient concerns
Why this verdict
The abstract-level profile supports the general design—five LLMs evaluated on multiple-choice breast cancer questions, clinical case analyses, and patient concerns—and identifies some models in results. However, the supplied abstract profile does not verify the full model list as stated, nor the exact counts of 90 multiple-choice questions, 10 clinical cases, and 20 patient concerns. These specifics require deeper-than-abstract evidence.
Study evidence
Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
Claim 3 of 6Not coveredHuman experts gave ChatGPT-5.2 and DeepSeek the highest completeness scores, while Gemini 3.0 received the highest expert-rated readability score; DeepSeek also had the lowest reading-difficulty score in the automated Chinese-language assessment.View evidenceHide evidence
As statedhighest completeness; highest readability; lowest reading-difficulty score
Why this verdict
The abstract-level profile supports that Gemini 3.0 had the highest reported expert-rated readability score and that DeepSeek had a lower/more stable automated reading-difficulty profile. It also supports that ChatGPT-5.2 and DeepSeek had relatively high completeness scores. However, the story’s stronger statement that ChatGPT-5.2 and DeepSeek received the highest completeness scores is not fully verifiable from the abstract profile because full domain-specific rankings and complete LMM outputs are not supplied.
Study evidence
Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
Study evidence
Automated LDU-TGP analysis showed that LLM-generated breast cancer texts had varying linguistic complexities across models.
“Text readability was analyzed using LDU-TGP.”
Claim 4 of 6Not coveredThe article says the study did not test patient comprehension or the safety and effectiveness of these models in direct patient-world clinical decision-making, and that all 10 clinical cases came from a single hospital.View evidenceHide evidence
As stated10 clinical cases from a single hospital
Why this verdict
The supplied abstract-level paper profile does not report the limitations that the study did not test patient comprehension, did not test safety/effectiveness in direct real-world clinical decision-making, or that all 10 clinical cases came from a single hospital. These may appear in the full paper, but they are not verifiable from the supplied abstract-depth profile.
Claim 5 of 6SupportedA comparative study in Scientific Reports evaluated modern large language models for providing breast cancer health information.View evidenceHide evidence
Why this verdict
The abstract-level profile supports that the paper was a comparative evaluation of five LLMs answering breast cancer–related questions across multiple task contexts, with specialist ratings and statistical comparison. The publication venue is not central to the scientific claim and is consistent with the supplied paper identifier/profile context.
Study evidence
Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
Claim 6 of 6SupportedAll five models showed high accuracy on standardized breast cancer knowledge questions, ranging from 82.22% to 94.44%.View evidenceHide evidence
As stated82.22% to 94.44%
Why this verdict
The paper profile explicitly reports multiple-choice test accuracy across models ranging from 82.22% to 94.44%, with broadly similar performance on standardized breast cancer knowledge questions.
Study evidence
Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
Context layer
What the story carried across
Nothing material from the study was dropped.
5 things the story did carry across
- The paper’s central design was a comparative in-silico evaluation of five LLMs on breast cancer multiple-choice knowledge questions, clinical case analyses, and patient concerns, with blinded specialist ratings across completeness, correctness/accuracy, readability, helpfulness, and safety.
- Multiple-choice breast cancer knowledge accuracy ranged from 82.22% to 94.44%, and the abstract profile reports no significant pairwise differences for this standardized knowledge test performance.
- Expert-rated multi-domain quality differed across models; ChatGPT-5.2 and DeepSeek had relatively high completeness scores, and Gemini 3.0 had the highest reported readability score.
- Automated LDU-TGP readability/linguistic-complexity analysis found model differences, including higher reading difficulty for ChatGPT-5.2 and lower/more stable reading difficulty for DeepSeek.
- The paper’s high-level conclusion was that no single LLM dominated across all assessed domains, supporting evaluation across multiple quality dimensions rather than relying on one aggregate winner.
Study layer
Study at a glance
Scan the study first. Expand only the parts you want to inspect.
Pieces of work
3
Evidence read
study summary
Lead result
in silico
1Lead resultin silicoBenchmark and compare multiple large language models (LLMs) on breast cancer–related question answering across several task contexts (multiple-choice knowledge questions, clinical case analyses, and patient concerns) using blinded specialist ratings and statistical comparison.Comparative LLM evaluation (in silico)ExpandCollapse
In plain English
Comparative evaluation of five large language models (LLMs) answering breast cancer–related items across three task contexts (multiple-choice knowledge questions, clinical case analyses, patient concerns). Responses were rated blindly by three specialist raters on completeness, correctness/accuracy, readability, helpfulness, and safety; text-level readability/linguistic complexity was additionally measured with the LDU-TGP platform. Statistical comparisons used Cochran’s Q, McNemar’s test, intraclass correlation coefficient (ICC), and linear mixed-effects models (LMM) with Bonferroni correction for multiple comparisons.
Key findings
- Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
- Linguistic/readability analysis (LDU-TGP) showed variation among models: ChatGPT-5.2 had higher reading difficulty scores while DeepSeek had lower and more stable scores.
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
2in silicoQuantify and compare readability/linguistic complexity of LLM-generated breast cancer responses using an automated readability/linguistic complexity platform (LDU-TGP).in silico automated readability analysis (LDU-TGP)ExpandCollapse
In plain English
Automated readability and linguistic complexity analysis of LLM-generated breast cancer texts using the LDU-TGP platform. The platform identified model-specific differences in linguistic complexity: ChatGPT-5.2 had a higher reading difficulty score, while DeepSeek showed a lower and more stable reading-difficulty profile; overall LLM-generated texts exhibited varying linguistic complexities across models.
Key findings
- Automated LDU-TGP analysis showed that LLM-generated breast cancer texts had varying linguistic complexities across models.
- ChatGPT-5.2 had a higher reading difficulty score, whereas DeepSeek had a lower and more stable reading difficulty score according to LDU-TGP.
“Text readability was analyzed using LDU-TGP.”
What this piece can’t prove
3 further details could not be confirmed from the summary.
3secondary dataAssess rater agreement/reliability (e.g., ICC) and characterize specialist concern/quality scoring behavior across many ratings.Inter-rater reliability analysis (ICC); descriptive summary of specialist ratingsExpandCollapse
In plain English
The abstract reports that three breast cancer specialists independently and blindly rated LLM-generated responses 1,500 times to assess sense of concern, yielding a mean score of 3.982 ± 0.460. Intraclass correlation coefficient (ICC) is listed among the statistical analyses, but no ICC estimate or details are reported in the abstract.
Key findings
- Three breast cancer experts independently assessed the sense of concern for patients 1,500 times and obtained a total mean score of 3.982 ± 0.460.
“Statistical analyses included ... Intraclass correlation coefficient (ICC)”
What this piece can’t prove
3 further details could not be confirmed from the summary.
Method layer
NewsLink found the paper. Tessa takes you deeper.
NewsLink checks the story. Tessa is where you inspect the paper, authors, evidence, and research context.
Open the paper in Tessa
A comparative study of large language models in responding to breast cancer–related questions
Scientific Reports · 2026
Why this one
Near certain
NewsLink found the paper. Tessa is where you inspect it deeply.
Papers considered
The selected paper, plus nearby candidates.
Crossref, PubMed, Europe PMC · 16 candidate papers
A comparative study of large language models in responding to breast cancer–related questions
Scientific Reports · 2026 · Crossref
Evaluating the Performance of Large Language Models for Breast Cancer Patient Education: A Comparative Study.
Journal of Cancer Education : the Official Journal of the American Association for Cancer Education · 2026 · PubMed, Europe PMC
Natural Language Question Answering on Tabular Data Using Large Language Models
Crossref
Development of a machine learning model for automatic data extraction from breast cancer pathology reports.
Scientific Reports · 2025 · PubMed
Reference Hallucination, Citation Reliability, and Readability of Large Language Models in Anatomy‐Related Question Answering
Clinical Anatomy · 2026 · Crossref
Can Small Open-Source Language Models With Retrieval-Augmented Generation Match GPT-4 Performance in Breast Cancer Clinical Decision Support?
2026 · Europe PMC
And 10 more candidates considered.