Clinical quality measurement commonly depends on manual abstraction of medical records, a process that is resource-intensive, burdensome, and often impractical for measures requiring interpretation of narrative text. These operational constraints have influenced which measures are developed and adopted, with some clinically important measures excluded because they are difficult to operationalize at scale. The authors evaluated whether a neuro-symbolic AI (NSAI) approach could reliably abstract complex quality measures from narrative pathology reports and thereby address these limitations.
The NSAI system evaluated in this work integrates two elements: large language model–based text extraction and symbolic reasoning. Each target quality measure is decomposed into atomic questions to structure the abstraction task. The system is aligned to real-world pathology reports through an iterative, human-in-the-loop method the authors call case-based refinement, in which prompts and symbolic rules are refined using real report examples.
Architectural decomposition—separating extraction and symbolic reasoning components—was used to stabilize performance across different underlying language-model backends. The authors report that this decomposition notably reduced performance variance across backends, a property considered important for clinical deployment.
The evaluation dataset comprised 2,000 pathology reports that had been independently double-abstracted. Trained human abstractors produced labels that were adjudicated to form the gold-standard reference used for comparison. The measures assessed were four pathology quality measures established by the College of American Pathologists, including Gastrointestinal Metaplasia (CAP 43).
When compared with the adjudicated gold standard, the NSAI system achieved high agreement: Cohen’s kappa = 0.95. Trained human abstractors measured against the same adjudicated standard reached Cohen’s kappa = 0.92. The NSAI system matched or modestly exceeded the human abstractors’ agreement overall, and it performed particularly well on the Gastrointestinal Metaplasia measure (CAP 43).
The authors performed component analyses to identify which elements of the NSAI pipeline drove performance gains:
Case-based refinement had the largest impact on accuracy, producing kappa improvements of up to 0.25. This shows iterative alignment to real examples was critical for reliable abstraction from variable narrative documentation.
Architectural decomposition mainly reduced performance variance across language-model backends by more than tenfold, improving consistency and repeatability—features important for clinical implementation where different model backends might be used.
These findings indicate both alignment to real reports and system architecture are important: refinement drives accuracy, while decomposition stabilizes results across technical variations.
The authors interpret their results to suggest that automated abstraction with NSAI could enable census-level quality measurement, lower the reporting burden on human abstractors, and expand the set of clinically meaningful measures that can realistically be operationalized from narrative clinical documentation. By reliably extracting measure elements from free-text pathology reports, the approach could change which quality measures are feasible to implement at scale.
Ethical oversight: the WCG Institutional Review Board waived ethical approval for this work; the use of de-identified pathology reports for secondary research was determined exempt under 45 CFR 46.104(d)(4) (Protocol 20251925).
Competing interests: two authors (F.B. and A.C.) are founders of and hold equity in Pharos Health, which developed the NSAI system evaluated in this study; one author (L.T.) is an employee of Pharos Health. Two authors (C.S. and G.B.) declared no competing interests.
Data availability: the pathology reports analyzed contain protected health information and are not publicly available. The authors state that the code and measure prompts that support the findings are available from the corresponding author upon reasonable request.
This article is a preprint and has not been peer reviewed; the authors and medRxiv note it reports new medical research that has not been certified by peer review and should not be used to guide clinical practice. The source does not provide additional external validation cohorts or implementation outcomes beyond the 2,000-report evaluation described.
A neuro-symbolic approach combining large language model extraction, symbolic reasoning, and iterative case-based refinement achieved high agreement with an adjudicated gold standard when abstracting pathology quality measures from narrative reports. The NSAI system matched or modestly exceeded trained human abstractors in overall agreement (kappa 0.95 vs 0.92) and showed that refinement and architectural decomposition respectively drove accuracy gains and variance reduction. The authors propose that automated abstraction could expand and scale quality measurement derived from narrative clinical documentation.