This retrospective single-centre study aimed to evaluate implementation characteristics of a commercially available chest radiography AI system within a large, real-world health check-up cohort. The investigators assessed the system at the workflow level, using lesion-specific metrics, and performed exploratory retrospective timeline analyses for histopathologically confirmed lung cancer cases. Because routine radiologist interpretation rather than universal CT or pathology served as the operational reference, primary analyses are presented as radiologist-referenced concordance evaluations rather than diagnostic accuracy against a gold-standard for all cases.
The dataset comprised 298,991 consecutive health check-up chest radiographs obtained between 2019 and 2023 from 114,866 individuals. All radiographs were interpreted under routine double reading by board-certified radiologists at the single centre. The study used routine radiologist judgments as the reference framework for operational concordance analyses; universal confirmatory imaging or histopathologic verification was not available for the entire cohort. A subset of cases (48) had histopathologic confirmation of lung cancer and was used in exploratory timeline analyses.
A commercially available AI system was evaluated using two prespecified approaches. The first, termed the all-score analysis, classified an image as AI-positive if any of ten AI-reported findings exceeded the manufacturer-recommended score threshold of 15. The second, the nodule-focused analysis, considered positivity only when the AI reported a nodule or mass at the same threshold. Both analyses applied the manufacturer-recommended score cutoff of 15 for AI positivity.
When compared to routine radiologist judgement, the AI system demonstrated the following performance metrics.
In the all-score analysis, radiologist-referenced sensitivity was 72.0% and specificity was 79.6%. The negative predictive value in this analysis was 99.0%.
In the nodule-focused analysis, radiologist-referenced sensitivity increased to 87.1% with specificity of 91.8%. The negative predictive value in the nodule-focused analysis was reported as 100.0%.
These results indicate stable concordance of the AI with routine radiologist readings across a very large health-check dataset, and particularly high sensitivity and negative predictive value for radiologist-reported pulmonary nodules and masses under the nodule-focused definition.
Among 48 histopathologically confirmed lung cancer cases identified within the cohort, retrospective timeline analyses were performed to compare the timing of AI positivity against routine radiologist positivity. In a subset of these cases, the AI flagged abnormalities earlier than routine radiologist readings. The authors explicitly designate these observations as exploratory and caution that they do not establish prospective clinical benefit or infer improved patient outcomes; they indicate only that earlier AI positivity occurred in some retrospective timelines within this dataset.
The study received institutional review board approval from the Jikei University School of Medicine (approval number 35-386 [12023]). Because this was a retrospective analysis using health check-up chest X-rays, AI outputs, and clinical information from the facility, the IRB approved a waiver of written informed consent and an opt-out procedure was used. The authors declared no competing interests. The manuscript states that data produced are available upon reasonable request to the authors.
In this large health check-up cohort, the evaluated chest radiography AI showed consistent operational concordance with routine radiologist judgement and high sensitivity for radiologist-reported pulmonary nodules and masses, particularly under the nodule-focused analysis using the manufacturer-recommended threshold of 15. High negative predictive values were reported in both prespecified analyses.
Interpretation is constrained by the study design: routine radiologist interpretation, rather than universal CT or pathological confirmation, served as the reference approach for primary analyses, so metrics reflect concordance with clinical workflow rather than independent diagnostic verification for all images. The exploratory finding that AI flagged some histopathologically confirmed lung cancers earlier than routine radiology in a subset of cases is hypothesis-generating and requires prospective evaluation to determine clinical impact.
Overall, the authors conclude that these operational concordance results and exploratory timelines support further prospective study of AI-assisted chest radiography workflows in health check-up settings, while acknowledging that the present findings do not establish prospective benefit or guide clinical practice.