Real-world oncology data are critical for clinical research and for delivering precision oncology care. The study addressed a common barrier: many genomic biomarker results are embedded in scanned, unstructured clinical documents and therefore are not immediately available for registries or secondary use. Manual abstraction of these documents delays availability of genomic data and slows real-world evidence generation. The investigation evaluated whether automated approaches using open-source optical character recognition (OCR) could accurately extract Oncotype DX recurrence scores from scanned reports, improving timeliness and reducing manual effort.
The investigators used a corpus of 675 Oncotype DX reports obtained from a Midwestern U.S. health system. Three OCR approaches were implemented and compared: Tesseract, EasyOCR, and a hybrid OCR implementation (details of the hybrid architecture were reported in the source). Extracted recurrence score values produced by each OCR approach were compared with values obtained through manual abstraction and with the local cancer registry’s recorded scores.
Performance evaluation included standard information-extraction metrics (agreement with manual abstraction, precision, recall, and F1 score) and measurement of processing time to characterize efficiency gains. In addition, multivariable logistic regression was applied to identify factors associated with discordance between registry-reported scores and manually abstracted scores.
Among the three OCR approaches, the hybrid OCR approach achieved the highest performance versus manual abstraction. Reported performance metrics for the hybrid approach included 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. These values indicate that the hybrid system produced very high accuracy in extracting Oncotype DX recurrence scores from scanned reports in this dataset.
The study noted that registry abstraction showed comparable performance to OCR-derived extraction but required substantially more manual effort. The source did not report specific error types, per-document failure modes, or comparisons of per-tool error distributions beyond the summarized metrics above.
Automated extraction using OCR substantially reduced processing time compared with manual abstraction, while maintaining high accuracy. The source highlights the potential of automated extraction to accelerate capture of genomic information into registries and other real-world data resources. Exact processing-time numbers and relative time-savings were not reported in the abstract; the conclusion emphasizes reduced latency as a principal benefit of automation.
A multivariable logistic regression was conducted to investigate predictors of discordance between registry-reported and manually abstracted Oncotype DX scores. According to the source, registry discordance was largely independent of examined patient and tumor characteristics. The only significant predictor identified was unknown progesterone receptor (PR) status. The abstract did not provide additional covariates, effect sizes, confidence intervals, or p-values in the summary; these details were not reported in the provided source text.
The authors conclude that automated extraction of genomic biomarkers from scanned clinical documents is a viable, scalable strategy to reduce delays in cancer data availability. Implementing accurate OCR-based extraction—particularly the hybrid approach evaluated—can preserve high data quality while reducing manual labor and processing time. Earlier capture of genomic results in registries and data warehouses may support modernization efforts for cancer registries and enhance real-world evidence generation in precision oncology.
The publication lists keywords including Breast cancer, Genomics, Oncology, and Optical character recognition. The authors declared no competing interests in the conflict of interest statement.
Notes on source limitations and unreported details
This rewrite is based solely on the source abstract and accompanying PubMed metadata. The abstract supplies overall performance metrics and high-level findings but does not include granular details such as the hybrid method’s technical architecture, per-tool failure modes, sample selection criteria, exact processing-time measurements, or full regression statistics. Those details were not reported in the provided source excerpt and would require consulting the full text of the article for deeper methodological and statistical information.