Skip to main content
Tessa NewsLink
Paste a health news link, or browse

Source study found

Story checked

Breast cancer AI test finds no single chatbot excels across every measure (opens in a new tab)

news-medical.net · 2026-09-10

Short answerEvidenceSource

Short answer

Mixed

Mixed.

The claims we could check match the study, but some claims were not covered by the evidence reviewed.

  • 2 supported
  • 4 not covered

Checked against the study summary. The full text wasn't available, so some details couldn't be settled either way.

Share this check

Follow the evidence trail
1
2

NewsLink checks it

Mixed

Every claim we could check holds up. Two of six claims match the study. This overall rating is based only on the claims we could check. Four claims the study doesn't address.

  • 2 supported
  • 4 not covered
Open claim evidence
3
Then inspect each claim

Evidence layer

Claim by claim

Each claim gets a verdict. Expand it to see the evidence directly below.

6 claims in this story

Showing all 6 claimsChoose a verdict to focus the list.

Then look for missing context

Context layer

What the story carried across

Nothing material from the study was dropped.

5 things the story did carry across
  • The paper’s central design was a comparative in-silico evaluation of five LLMs on breast cancer multiple-choice knowledge questions, clinical case analyses, and patient concerns, with blinded specialist ratings across completeness, correctness/accuracy, readability, helpfulness, and safety.
  • Multiple-choice breast cancer knowledge accuracy ranged from 82.22% to 94.44%, and the abstract profile reports no significant pairwise differences for this standardized knowledge test performance.
  • Expert-rated multi-domain quality differed across models; ChatGPT-5.2 and DeepSeek had relatively high completeness scores, and Gemini 3.0 had the highest reported readability score.
  • Automated LDU-TGP readability/linguistic-complexity analysis found model differences, including higher reading difficulty for ChatGPT-5.2 and lower/more stable reading difficulty for DeepSeek.
  • The paper’s high-level conclusion was that no single LLM dominated across all assessed domains, supporting evaluation across multiple quality dimensions rather than relying on one aggregate winner.
Then read the study layer

Study layer

Study at a glance

Scan the study first. Expand only the parts you want to inspect.

Pieces of work

3

Evidence read

study summary

Lead result

in silico

1Lead resultin silicoBenchmark and compare multiple large language models (LLMs) on breast cancer–related question answering across several task contexts (multiple-choice knowledge questions, clinical case analyses, and patient concerns) using blinded specialist ratings and statistical comparison.Comparative LLM evaluation (in silico)Expand

In plain English

Comparative evaluation of five large language models (LLMs) answering breast cancer–related items across three task contexts (multiple-choice knowledge questions, clinical case analyses, patient concerns). Responses were rated blindly by three specialist raters on completeness, correctness/accuracy, readability, helpfulness, and safety; text-level readability/linguistic complexity was additionally measured with the LDU-TGP platform. Statistical comparisons used Cochran’s Q, McNemar’s test, intraclass correlation coefficient (ICC), and linear mixed-effects models (LMM) with Bonferroni correction for multiple comparisons.

Key findings

  • Multiple-choice accuracy across five evaluated LLMs ranged from 82.22% to 94.44%, with no significant pairwise differences reported.82.22%–94.44% accuracy; no significant pairwise differences
  • Linguistic/readability analysis (LDU-TGP) showed variation among models: ChatGPT-5.2 had higher reading difficulty scores while DeepSeek had lower and more stable scores.
“Five LLMs were evaluated on breast cancer multiple-choice questions, clinical case analyses, and patient concerns.”
2in silicoQuantify and compare readability/linguistic complexity of LLM-generated breast cancer responses using an automated readability/linguistic complexity platform (LDU-TGP).in silico automated readability analysis (LDU-TGP)Expand

In plain English

Automated readability and linguistic complexity analysis of LLM-generated breast cancer texts using the LDU-TGP platform. The platform identified model-specific differences in linguistic complexity: ChatGPT-5.2 had a higher reading difficulty score, while DeepSeek showed a lower and more stable reading-difficulty profile; overall LLM-generated texts exhibited varying linguistic complexities across models.

Key findings

  • Automated LDU-TGP analysis showed that LLM-generated breast cancer texts had varying linguistic complexities across models.
  • ChatGPT-5.2 had a higher reading difficulty score, whereas DeepSeek had a lower and more stable reading difficulty score according to LDU-TGP.
“Text readability was analyzed using LDU-TGP.”
What this piece can’t prove

3 further details could not be confirmed from the summary.

3secondary dataAssess rater agreement/reliability (e.g., ICC) and characterize specialist concern/quality scoring behavior across many ratings.Inter-rater reliability analysis (ICC); descriptive summary of specialist ratingsExpand

In plain English

The abstract reports that three breast cancer specialists independently and blindly rated LLM-generated responses 1,500 times to assess sense of concern, yielding a mean score of 3.982 ± 0.460. Intraclass correlation coefficient (ICC) is listed among the statistical analyses, but no ICC estimate or details are reported in the abstract.

Key findings

  • Three breast cancer experts independently assessed the sense of concern for patients 1,500 times and obtained a total mean score of 3.982 ± 0.460.
“Statistical analyses included ... Intraclass correlation coefficient (ICC)”
What this piece can’t prove

3 further details could not be confirmed from the summary.

Finally, the search trail

Method layer

NewsLink found the paper. Tessa takes you deeper.

NewsLink checks the story. Tessa is where you inspect the paper, authors, evidence, and research context.

Open the paper in Tessa

A comparative study of large language models in responding to breast cancer–related questions

Scientific Reports · 2026

Why this one

Near certain

NewsLink found the paper. Tessa is where you inspect it deeply.

Papers considered

The selected paper, plus nearby candidates.

Crossref, PubMed, Europe PMC · 16 candidate papers

Selected

A comparative study of large language models in responding to breast cancer–related questions

Scientific Reports · 2026 · Crossref

Candidate

Evaluating the Performance of Large Language Models for Breast Cancer Patient Education: A Comparative Study.

Journal of Cancer Education : the Official Journal of the American Association for Cancer Education · 2026 · PubMed, Europe PMC

And 10 more candidates considered.