Source study found
Story checked
AI reads doctors' notes at scale, revealing data absent from coded medical records (opens in a new tab)
medicalxpress.com · 2026-09-18
Short answer
Mostly not supportedMostly not supported.
The claims we could check match the study, but some claims were not covered by the evidence reviewed.
- 1 supported
- 6 not covered
Checked against the study summary. The full text wasn't available, so some details couldn't be settled either way.
Share this check
The story
AI reads doctors' notes at scale, revealing data absent from coded medical records
medicalxpress.com · 2026-09-18
The story’s checkable claims.
Read the original story (opens in a new tab)NewsLink checks it
Mostly not supported
The one claim we could check holds up. One of seven claims matches the study. This overall rating is based only on the claims we could check. Six claims the study doesn't address.
- 1 supported
- 6 not covered
The source study
Computable longitudinal patient journeys from structured and unstructured EHR data
Evidence layer
Claim by claim
Each claim gets a verdict. Expand it to see the evidence directly below.
Reading mode
Scan verdicts. Open evidence only when needed.
Browse by verdict
7 claims in this storyShowing all 7 claimsChoose a verdict to focus the list.
Claim 1 of 7Not coveredResearchers described an AI system that accurately reads clinicians’ written notes at scale, converts extracted facts into analyzable data, and links every fact back to the exact sentence in the chart.View evidenceHide evidence
Why this verdict
The abstract-level profile supports that the paper presents a large-pretrained-language-model framework to extract computable clinical data from unstructured EHR text, describes it as scalable, and reports high physician-adjudicated accuracy. However, the abstract-level evidence does not verify the story's stronger implementation detail that every extracted fact is linked back to the exact source sentence in the chart.
Study evidence
Large, pre-trained language models can be used to extract computable clinical data from unstructured EHR text.
“we present a novel approach that uses large, pre-trained language models to accurately extract computable clinical data from unstructured EHR text”
Study evidence
LLM-extracted clinical entities demonstrated high accuracy when compared with a blinded expert reference standard, as confirmed by physician adjudication.
“Physician adjudication confirmed high accuracy against a blinded expert reference standard, with high expert physician inter-reviewer agreement.”
Claim 2 of 7Not coveredBoard-certified physicians independently reviewed the system’s output against original notes, and the system scored 99.4% on a combined measure of correctness and completeness compared with an adjudicated standard.View evidenceHide evidence
As stated99.4%
Why this verdict
The abstract-level profile supports physician adjudication against a blinded expert reference standard and high inter-reviewer agreement, but it does not provide the reported 99.4% combined correctness/completeness metric, confirm board certification, or give enough detail on independent review against original notes. The numerical claim is therefore not verifiable at abstract depth.
Study evidence
LLM-extracted clinical entities demonstrated high accuracy when compared with a blinded expert reference standard, as confirmed by physician adjudication.
“Physician adjudication confirmed high accuracy against a blinded expert reference standard, with high expert physician inter-reviewer agreement.”
Claim 3 of 7Not coveredThe article says the team applied the approach to more than 16,000 adults who started a GLP-1 medication in routine outpatient care, including people using it for diabetes and for weight loss.View evidenceHide evidence
As statedmore than 16,000 adults
Why this verdict
The profile supports that the framework was demonstrated in a large-scale longitudinal analysis of individuals initiating GLP-1 receptor agonists. It does not verify the cohort size of more than 16,000 adults, the routine outpatient-care framing, or the indication split between diabetes and weight loss at abstract depth.
Study evidence
Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
Claim 4 of 7Not coveredAmong GLP-1 users, people with healthier blood sugar at baseline lost more weight and reached 5% weight loss faster, while people with the worst baseline glucose control improved blood sugar the most and fastest.View evidenceHide evidence
As stated7.7% vs 2.7% body-weight loss at 12 months; about 210 vs 413 days to 5% weight loss
Why this verdict
The abstract-level profile supports longitudinal modeling of weight and HbA1c after GLP-1 RA initiation and associations with treatment response. It does not report the baseline-glycemia subgroup pattern, the 12-month weight-loss percentages, or the time-to-5%-weight-loss estimates, so the specific findings are not verifiable at this depth.
Study evidence
Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
Claim 5 of 7Not coveredThe article says weight-loss outcomes also differed by sex and age, with women and adults ages 20 to 39 losing more body weight than men and older adults.View evidenceHide evidence
As stated6.1% vs 4.0%; 8.1% vs 5.1%
Why this verdict
The profile supports a GLP-1 RA longitudinal outcomes analysis but does not mention sex- or age-stratified weight-loss outcomes or the specific percentages reported by the story. This demographic-modifier claim is not verifiable from the abstract-level evidence.
Study evidence
Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
Claim 6 of 7Not coveredThe story says the note-reading method surfaced outcomes rarely available in coded data, including improved depression scores, reduced pain intensity, and decreased waist circumference among subsets of patients.View evidenceHide evidence
As statedabout six-point drop in depression scores; about 2.7 inches (6.9 cm) waist reduction
Why this verdict
The abstract-level profile supports that the method characterized outcomes captured only in unstructured clinical notes and enabled time-to-event analyses. It does not identify depression scores, pain intensity, or waist circumference, nor does it provide the reported magnitude of depression-score or waist-circumference changes. The general note-derived-outcome concept is supported, but the specific examples and estimates are not verifiable at abstract depth.
Study evidence
Large, pre-trained language models can be used to extract computable clinical data from unstructured EHR text.
“we present a novel approach that uses large, pre-trained language models to accurately extract computable clinical data from unstructured EHR text”
Study evidence
Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
Claim 7 of 7SupportedThe article emphasizes that the work was observational rather than a controlled trial and that the findings cannot by themselves establish cause and effect.View evidenceHide evidence
Why this verdict
The profile characterizes the GLP-1 analysis as a retrospective, secondary-data EHR cohort/observational study and notes the absence of abstract-level details on confounding control. The story's caveat that the work was observational rather than a controlled trial and cannot by itself establish cause and effect is consistent with that design.
Study evidence
Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
Context layer
What the story left out
Important study details the story did not include.
Integration of extracted entities with structured EHR data and medical ontologies into a knowledge graph, with an agent-friendly programmatic query interface.
The paper profile treats knowledge-graph construction, ontology embedding, and an agent-friendly query interface as a material methodological component. The supplied story presentation focuses on note extraction and computable data but does not reflect the KG/ontology/query-interface element.
From Knowledge graph construction, ontology embedding, and programmatic query interface
6 things the story did carry across
- Disease-agnostic LLM framework for extracting computable clinical entities from unstructured EHR notes at scale.
- Physician-adjudicated validation against a blinded expert reference standard with high accuracy and high inter-reviewer agreement.
- Large-scale retrospective longitudinal analysis of GLP-1 receptor agonist initiators using integrated structured and note-derived EHR data.
- Modeling of longitudinal changes in weight and HbA1c and associations with GLP-1 treatment response.
- Characterization of outcomes captured only in unstructured clinical notes, including time-to-event outcomes.
- Observational, retrospective EHR design and resulting limits on causal inference for GLP-1 outcome comparisons.
Study layer
Study at a glance
Scan the study first. Expand only the parts you want to inspect.
Pieces of work
4
Evidence read
study summary
Lead result
secondary data
1Lead resultsecondary dataDevelop a disease-agnostic framework using large pre-trained language models to extract computable clinical entities from unstructured EHR text and integrate them with structured EHR data into an ontology-embedded knowledge graph with an agent-friendly query interface to enable scalable real-world evidence generation.secondary data, computational extraction and integration pipelineExpandCollapse
In plain English
The paper introduces a disease-agnostic computational pipeline that uses large, pre-trained language models to extract computable clinical entities from unstructured EHR text, integrates these entities with structured EHR data and medical ontologies into a patient-level knowledge graph, and exposes the KG via an agent-friendly programmatic query interface to support scalable real-world evidence generation. The authors report physician adjudication against a blinded expert reference standard with high accuracy and high inter-reviewer agreement, and demonstrate the framework by reconstructing longitudinal patient trajectories and modeling weight and HbA1c changes after initiation of GLP-1 receptor agonist therapy.
Key findings
- Large, pre-trained language models can be used to extract computable clinical data from unstructured EHR text.
- Extracted entities can be integrated with structured EHR data and medical ontologies and organized into a knowledge graph enabling interrogation of relationships at patient and cohort levels.
“we present a novel approach that uses large, pre-trained language models to accurately extract computable clinical data from unstructured EHR text”
What this piece can’t prove
- Abstract does not provide numeric performance metrics (e.g., precision, recall) or sample sizes for the extraction validation.
3 further details could not be confirmed from the summary.
2secondary dataDevelop a disease-agnostic framework using large pre-trained language models to extract computable clinical entities from unstructured EHR text and integrate them with structured EHR data into an ontology-embedded knowledge graph with an agent-friendly query interface to enable scalable real-world evidence generation.Knowledge graph construction, ontology embedding, and programmatic query interfaceExpandCollapse
In plain English
Authors describe a disease-agnostic framework that integrates clinical entities extracted from unstructured EHR text with structured EHR fields, embeds those entities using medical ontologies, and organizes the integrated data into a knowledge graph (KG) accessible via an agent-friendly programmatic query interface to support patient-level and population-scale interrogation and downstream analyses.
Key findings
- Extracted clinical entities were integrated with structured EHR data, embedded with medical ontologies, and organized into a knowledge graph.
- The knowledge graph enables interrogation of relationships between variables at the level of individual patients and at scale across patient populations.
“The extracted clinical entities are integrated with structured EHR data, embedded with medical ontologies, and organized into a knowledge graph (KG)”
What this piece can’t prove
4 further details could not be confirmed from the summary.
3secondary dataValidate the accuracy of LLM-extracted clinical data from unstructured EHR notes against a blinded expert reference standard with physician adjudication and inter-reviewer agreement.Expert chart review / validationExpandCollapse
In plain English
The authors validated LLM-extracted clinical entities from unstructured EHR notes by comparing extractions to a blinded expert reference standard using physician adjudication; adjudication confirmed high accuracy and showed high inter-reviewer agreement.
Key findings
- LLM-extracted clinical entities demonstrated high accuracy when compared with a blinded expert reference standard, as confirmed by physician adjudication.
- Expert physician adjudicators showed high inter-reviewer agreement in the validation process.
“Physician adjudication confirmed high accuracy against a blinded expert reference standard, with high expert physician inter-reviewer agreement.”
What this piece can’t prove
3 further details could not be confirmed from the summary.
4secondary dataDemonstrate clinical utility by constructing computable longitudinal patient journeys and performing a large-scale longitudinal real-world analysis of outcomes following GLP-1 receptor agonist initiation, including continuous outcomes (weight, HbA1c) and time-to-event outcomes captured in unstructured notes.Retrospective EHR cohort with NLP-derived variablesExpandCollapse
In plain English
Retrospective, large-scale longitudinal analysis of individuals initiating GLP-1 receptor agonists using integrated structured EHR data and NLP-extracted entities from clinical notes. The study used large pre-trained language models to extract clinical entities, integrated these with structured fields into a knowledge graph, and used this computable longitudinal representation to identify GLP-1 RA initiators, reconstruct patient-level trajectories, and model longitudinal changes in weight and HbA1c as well as time-to-event outcomes captured in unstructured notes. Physician adjudication against a blinded expert reference standard indicated high accuracy and high inter-reviewer agreement. The platform and KG enabled scalable, investigator-controlled analyses.
Key findings
- Pre-trained language-model-based extraction of clinical entities from unstructured EHR notes showed high accuracy when adjudicated by physicians against a blinded expert reference standard, with high inter-reviewer agreement.
- The integrated approach identified large numbers of patients initiating GLP-1 receptor agonists and reconstructed patient-level longitudinal trajectories.
“To demonstrate clinical utility, we conducted a large-scale longitudinal analysis of treatment responses among individuals initiating glucagon-like peptide-1 receptor agonists (GLP-1 RAs) therapy.”
What this piece can’t prove
- Abstract does not report cohort size, inclusion/exclusion criteria, baseline characteristics, or how confounding was addressed.
3 further details could not be confirmed from the summary.
Method layer
NewsLink found the paper. Tessa takes you deeper.
NewsLink checks the story. Tessa is where you inspect the paper, authors, evidence, and research context.
Open the paper in Tessa
Computable longitudinal patient journeys from structured and unstructured EHR data
Nature medicine · 2026
Why this one
Near certain
NewsLink found the paper. Tessa is where you inspect it deeply.
Papers considered
The selected paper, plus nearby candidates.
PubMed, Crossref, Europe PMC · 40 candidate papers
Computable longitudinal patient journeys from structured and unstructured EHR data
Nature Medicine · 2026 · PubMed, Crossref
Author Correction: When membrane proteins prefer lipids
Nature Chemical Biology · 2026 · Crossref
Intravascular Radiation Delays Recurrence of Recalcitrant Coronary In-Stent Restenosis.
2026 · Europe PMC
Author Correction: Advances, challenges and prospects of origami and kirigami optoelectronics
Nature Communications · 2026 · Crossref
Associations between genetic ancestry and allergic outcomes at 10 years among Black children.
2026 · Europe PMC
Author Correction: Broadening the molecular glue landscape
Nature Chemical Biology · 2026 · Crossref
And 34 more candidates considered.