Source study found
Story checked
AI tool successfully predicts outcomes of immunotherapy in lung cancer (opens in a new tab)
medicalxpress.com · 2026-09-14
Short answer
Mostly not supportedMostly not supported.
2 claims go further than the study. 3 other points were not covered by the paper.
- 1 supported
- 2 overstated
- 3 not covered
Checked against the study summary. The full text wasn't available, so some details couldn't be settled either way.
Share this check
The story
AI tool successfully predicts outcomes of immunotherapy in lung cancer
medicalxpress.com · 2026-09-14
The story’s checkable claims.
Read the original story (opens in a new tab)NewsLink checks it
Mostly not supported
Two of six claims overstate the study. One of six checks out. Three claims the study doesn't address.
- 1 supported
- 2 overstated
- 3 not covered
The source study
Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC
Evidence layer
Claim by claim
Each claim gets a verdict. Expand it to see the evidence directly below.
Reading mode
Scan verdicts. Open evidence only when needed.
Browse by verdict
6 claims in this storyShowing all 6 claimsChoose a verdict to focus the list.
Claim 1 of 6OverstatedThe study reported findings from the I3LUNG project, described as a large international trial aimed at improving treatment of metastatic NSCLC by developing AI-based predictive models to help determine the best therapeutic approach for each patient.View evidenceHide evidence
Why this verdict
The profile supports that I3LUNG is a large international real-world AI-based NSCLC study developing predictive models. However, the story’s framing as a large international “trial” aimed at determining the “best therapeutic approach” for each patient goes beyond the abstract-level evidence, which describes observational model development/validation and ongoing prospective validation rather than demonstrated treatment optimization or individualized therapeutic selection.
Study evidence
ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
Claim 2 of 6OverstatedThe researchers said their AI models outperformed standard clinical biomarkers, with one model using clinical and blood data reaching an AUC of 0.77 and a model using clinical, blood, imaging and digital pathology data reaching an AUC of 0.88.View evidenceHide evidence
As statedAUC 0.77 and 0.88
Why this verdict
The abstract profile supports that clinical+blood AI models reached AUC up to 0.77 in the independent TEST set and significantly outperformed PD-L1 and other standard biomarkers/scores in TEST. But the story’s multimodal AUC 0.88 is not present in the abstract profile, and the paper profile emphasizes that multimodal incremental benefit was uncertain and did not consistently translate to TEST or external validation. Presenting the multimodal result without that limitation overstates the evidence.
Study evidence
ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
Study evidence
CB-only ML and DL models achieved AUC up to 0.77 in the independent TEST set.AUC up to 0.77 (TEST)
“Although multimodal integration with MLEF (CB+CT+DP) was associated with higher performance, its incremental benefit remains uncertain, not translated in TEST and EXVAL.”
Claim 3 of 6Not coveredThe international research team enrolled 2,396 patients with advanced NSCLC treated with immunotherapy across six centers in Italy, Germany, Greece, Israel, Spain and the United States.View evidenceHide evidence
As stated2,396 patients; six centers; six countries
Why this verdict
The abstract profile supports the cohort size of 2,396 patients and an international real-world NSCLC immunotherapy cohort. It does not verify the stated six centers or the specific countries listed, so those detailed enrollment/location claims are not verifiable at abstract depth.
Study evidence
ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
Claim 4 of 6Not coveredIn a human-testing component, 20 physicians reviewed 100 real patient cases and access to the AI tool improved sensitivity for identifying responders, with AUC rising from 0.72 to 0.87.View evidenceHide evidence
As stated20 physicians; 100 cases; AUC 0.72 to 0.87
Why this verdict
The abstract profile supports the qualitative finding that lung expert and nonexpert physicians improved predictions using the XAI clinical+blood tool. It does not provide the claimed numbers of physicians, number of cases, sensitivity, or AUC change from 0.72 to 0.87, so the magnitude and design details are not verifiable at abstract depth.
Study evidence
Lung expert and nonexpert physicians improved their outcome predictions when assisted by the explainable AI (XAI) ML CB-only decision-support tool.
“The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool.”
Claim 5 of 6Not coveredThe article says the project’s retrospective phase is complete and that I3LUNG is now prospectively enrolling more than 2,000 patients across the same six centers to focus on treatment optimization.View evidenceHide evidence
As statedmore than 2,000 patients
Why this verdict
The abstract profile supports that prospective validation of the decision-support system is ongoing in more than 2,000 patients. It does not verify that the retrospective phase is complete, that enrollment is across the same six centers, or that the prospective phase is specifically focused on treatment optimization rather than validation, so the full claim is not verifiable at abstract depth.
Study evidence
ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
Study evidence
Lung expert and nonexpert physicians improved their outcome predictions when assisted by the explainable AI (XAI) ML CB-only decision-support tool.
“The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool.”
Claim 6 of 6SupportedA new study published in Nature Medicine found that AI tools can help physicians predict treatment and survival outcomes in patients with advanced non-small cell lung cancer treated with immunotherapy.View evidenceHide evidence
Why this verdict
At abstract level, the paper profile supports the broad claim that AI models were developed and validated to predict immunotherapy-related outcomes in NSCLC, and that physicians improved predictions when using an explainable AI tool. The abstract profile does not provide detailed outcome-by-outcome support for the exact phrase “treatment and survival outcomes,” but the overall framed claim is aligned with the paper’s stated contributions.
Study evidence
ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
Study evidence
Lung expert and nonexpert physicians improved their outcome predictions when assisted by the explainable AI (XAI) ML CB-only decision-support tool.
“The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool.”
Context layer
What the story left out
Important study details the story did not include.
AI models were evaluated in an independent TEST set and external validation cohorts, with external validation performance dropping to AUC 0.55–0.72.
The story mentions retrospective validation but does not report the interpretation-changing limitation that external validation performance declined substantially, indicating possible generalizability limits.
From observational cohort model development and validation; comparative multimodal fusion evaluation
Multimodal integration with clinical+blood, CT and digital pathology was associated with higher performance in development, but its incremental benefit was uncertain and did not translate consistently to TEST or external validation.
The story emphasizes a high multimodal AUC and superiority framing but does not mention the paper profile’s key limitation that multimodal incremental benefit was uncertain and not consistently validated.
From comparative multimodal fusion evaluation
5 things the story did carry across
- I3LUNG is a large international real-world observational cohort/model-development study enrolling 2,396 NSCLC patients to develop and validate AI models for immunotherapy-related outcome prediction.
- AI models significantly outperformed PD-L1, ECOG PS, NLR, LDH and LIPI in the independent TEST set.
- Clinical+blood-only ML/DL models achieved AUC up to 0.77 in the independent TEST set.
- A clinical usability study found that lung expert and nonexpert physicians improved predictions when using an explainable AI clinical+blood decision-support tool.
- Prospective validation of the AI decision-support system is ongoing in more than 2,000 patients.
Study layer
Study at a glance
Scan the study first. Expand only the parts you want to inspect.
Pieces of work
3
Evidence read
study summary
Lead result
secondary data
1Lead resultsecondary dataDevelop and validate real-world multimodal AI models (clinical+blood, CT imaging, digital pathology, genomics) to predict immunotherapy-related outcomes in NSCLC, including internal testing and external validation, and compare against standard biomarkers/scores (PD-L1, ECOG PS, NLR, LDH, LIPI).observational cohort model development and validationExpandCollapse
In plain English
I3LUNG (NCT05537922) is a large (n=2,396) international real-world cohort study that developed and evaluated multimodal AI models (clinical+blood, CT, digital pathology, genomics) to predict immunotherapy-related outcomes in NSCLC. The authors trained machine-learning (ML) and deep-learning (DL) models with early- and intermediate-fusion approaches, reported discrimination (AUC) in an independent internal TEST set and in external validation (EXVAL), benchmarked model performance against PD-L1 and clinical/hematologic prognostic scores, and performed a clinical usability assessment of an explainable AI decision-support tool.
Key findings
- ML and DL models using clinical and blood (CB-only) inputs achieved up to AUC = 0.77 in the independent TEST set.AUC up to 0.77
- Model discrimination decreased in external validation cohorts (EXVAL), with reported AUC range 0.55–0.72.AUC range 0.55–0.72
“I3LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients.”
What this piece can’t prove
- External validation performance declined (AUC 0.55–0.72), indicating possible population heterogeneity and limited generalizability.
- Incremental benefit of multimodal fusion is reported as uncertain and not consistently observed in TEST and EXVAL.
2 further details could not be confirmed from the summary.
2human in vivoEvaluate clinical usability of an explainable AI (XAI) decision-support tool by assessing whether expert and nonexpert physicians improve their outcome predictions when using the tool.clinical usability study (assisted vs unassisted physician predictions)ExpandCollapse
In plain English
A clinical usability study reported that both lung expert and nonexpert physicians improved their outcome predictions when using an explainable AI (XAI) decision-support tool based on a machine-learning model trained on clinical and blood (CB) data.
Key findings
- Lung expert and nonexpert physicians improved their outcome predictions when assisted by the explainable AI (XAI) ML CB-only decision-support tool.
“The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool.”
What this piece can’t prove
- No methodological details provided on study design (randomization, blinding, crossover, or parallel assignment).
- Unclear how clinician participants were recruited or characterized beyond 'expert' vs 'nonexpert'.
- No description of the prediction task, evaluation metrics, or whether training/practice with the tool occurred before measurement.
1 further detail could not be confirmed from the summary.
3secondary dataAssess the incremental benefit (or lack thereof) of multimodal integration (early fusion and intermediate fusion) versus CB-only models, including whether gains translate to independent test and external validation cohorts.comparative multimodal fusion evaluationExpandCollapse
In plain English
The abstract reports that integrating clinical+blood (CB) data with CT and digital pathology (DP) using an early-fusion machine-learning approach (MLEF) was associated with higher performance than CB-only models, but that this incremental benefit was uncertain and did not translate to the independent internal TEST set or to external validation (EXVAL). CB-only ML and DL models achieved AUC up to 0.77 in the TEST set, with lower AUCs in EXVAL (range 0.55–0.72). Details on effect sizes for the multimodal incremental gains, comparison metrics between MLEF and DLIF, and evaluation stratified by outcome are not provided in the abstract.
Key findings
- CB-only ML and DL models achieved AUC up to 0.77 in the independent TEST set.AUC up to 0.77 (TEST)
- Performance decreased in external validation (EXVAL), with AUCs reported in the range 0.55–0.72.AUC range 0.55–0.72 (EXVAL)
“Although multimodal integration with MLEF (CB+CT+DP) was associated with higher performance, its incremental benefit remains uncertain, not translated in TEST and EXVAL.”
What this piece can’t prove
- The abstract indicates non-translation of multimodal gains to TEST and EXVAL but does not specify whether this reflects overfitting, cohort heterogeneity, or other factors; causal reasons are not detailed.
2 further details could not be confirmed from the summary.
Method layer
NewsLink found the paper. Tessa takes you deeper.
NewsLink checks the story. Tessa is where you inspect the paper, authors, evidence, and research context.
Open the paper in Tessa
Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC
Nature medicine · 2026
Why this one
Near certain
NewsLink found the paper. Tessa is where you inspect it deeply.
Papers considered
The selected paper, plus nearby candidates.
PubMed, Crossref, Europe PMC · 39 candidate papers
Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC
Nature Medicine · 2026 · PubMed, Crossref
Author Correction: Evolution of claudin18.2 therapies in gastroesophageal cancers
Nature Medicine · 2026 · Crossref
What practicing pathologists and oncologists should know about the new computational pathology-based companion diagnostic tools.
2026 · Europe PMC
Artificial intelligence-generated synthetic data for cancer research and clinical trials
Nature Reviews Cancer · 2026 · Crossref
Network analysis predicts pembrolizumab response in advanced NSCLC with PD-L1 < 50.
2026 · Europe PMC
NATURE AND CULTURE: AN ECOLINGUISTICS’ ANALYSIS OF IS A RIVER ALIVE BY ROBERT MACFARLANE
Scholarly Journal · 2026 · Crossref
And 33 more candidates considered.