Automated surveillance of adverse events from free-text endoscopy procedure reports is an emerging tool for quality assurance and case-finding. The performance of such systems in non-English clinical text remains under-evaluated. This study compared a rule-based natural language processing approach with a transformer-based NLP model for classifying index procedure reports according to whether a clinician-adjudicated endoscopy-related adverse event (AE) occurred within 30 days.
The primary objective was to measure diagnostic accuracy of the two NLP methods for identifying attributable 30-day AEs using clinician adjudication as the reference standard.
This was a retrospective, single-centre diagnostic accuracy study conducted in a tertiary endoscopy unit in Türkiye. The review period spanned August 2015 to August 2025. The analysis focused on classifying individual index procedure reports using the report text alone, while clinician adjudication could use subsequent documentation within 30 days.
The source cohort comprised 140,385 endoscopy procedure reports. From this pool, all 1,512 reports identified by a pre-specified lexicon as potentially AE-positive (lexicon-positive) were selected for clinician review, together with a random sample of 2,000 lexicon-negative reports. Clinician adjudication of this set identified 1,208 AE-positive reports in the adjudicated subset.
A separate 13,333-report corpus—drawn and allocated at the patient level—was used for model development and evaluation. That corpus was split into training (n=9,333), validation (n=2,000) and a locked, fully adjudicated test set (n=2,000), in which 180 reports were AE-positive.
Two automated approaches were compared:
Both models made classifications using only the index procedure report; adjudication incorporated later documentation and clinical context to determine whether an AE was attributable within 30 days.
Model development used the training and validation splits. Performance metrics reported here are from the locked, fully adjudicated test set of 2,000 reports, which contained 180 AE-positive cases. Allocation was performed at the patient level to avoid information leakage between splits.
The primary outcome was confirmation of an attributable adverse event within 30 days of the index procedure, as determined by clinician adjudication. Secondary outcomes included AE category and timing of documentation, although specific breakdowns of AE categories and timing were not detailed in the source text.
Reported measures included sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score, and the transformer model’s precision–recall area under the curve (PR-AUC). Statistical comparison between models included McNemar’s test where applicable.
On the locked test set (n=2,000; 180 AE-positive):
The rule-based model achieved 84.4% sensitivity (95% CI 78.4% to 89.0%) and 99.5% specificity (95% CI 99.1% to 99.7%).
The transformer model achieved 92.2% sensitivity (95% CI 87.4% to 95.3%), a statistically significant improvement in sensitivity versus the rule-based model (McNemar’s test, p=0.01), and 98.4% specificity (95% CI 97.7% to 98.9%).
Positive predictive values were 94.4% for the rule-based model and 85.1% for the transformer. Negative predictive values were 98.5% and 99.2%, respectively.
Both models had identical F1 scores of 0.89 on the test set.
The transformer’s PR-AUC was reported as 0.91 (95% CI 0.87 to 0.94), summarizing its precision–recall performance across thresholds.
Compared with the rule-based approach, the transformer reduced the number of false negatives from 28 to 14, indicating fewer missed AEs. However, this increased the number of false positives from 9 to 29. In other words, the transformer traded improved sensitivity (fewer missed events) for reduced precision (more false alarms).
A key limitation noted in the source is that lexicon-negative reports were only partially verified by clinician adjudication. Because the lexicon-negative portion of the full 140,385-report cohort was not exhaustively adjudicated, the study did not estimate cohort-wide AE incidence. Other limitations inherent to single-centre retrospective designs and to using index-report–only input for model classification were acknowledged implicitly by the authors’ recommendations against autonomous use based solely on these results.
Both the rule-based and transformer models demonstrated high specificity, indicating strong ability to exclude non-AE reports in this test set. The transformer model provided higher sensitivity and a higher PR-AUC, reducing missed AEs but producing more false positives.
Given this trade-off, the authors conclude these automated methods are best suited to clinician-supervised retrospective case identification and quality-assurance workflows rather than to autonomous diagnosis, prospective prediction, or real-time point-of-care decision-making. Increased false positives with the transformer would require human review to avoid unnecessary actions.
The study authors state that external validation is required before broader adoption in other settings or languages.
In this ten-year, single-centre retrospective diagnostic accuracy study, both a rule-based and a transformer-based NLP approach achieved high specificity for detecting endoscopy-related AEs from non-English free-text reports. The transformer improved sensitivity and PR-AUC but generated more false positives. These results support clinician-supervised retrospective surveillance and quality assurance, but do not support autonomous clinical use without further validation and testing in external cohorts.