Retained bile duct stones can cause biliary obstruction, cholangitis, and pancreatitis. Interpretation of intraoperative cholangiography (IOC) can generate false positives that lead to additional procedures. This study evaluated whether deep-learning segmentation models can identify extrahepatic bile duct stones on representative IOC images at the case level and how completely individual stones are localized.
The primary objectives were to (1) quantify case-level diagnostic performance of segmentation-based deep-learning models against a composite clinical reference standard and (2) characterize completeness of individual-stone localization against expert-reviewed annotations.
Representative IOC images were retrospectively collected and annotated for extrahepatic biliary anatomy (including the common bile duct and common hepatic duct) and for stones. A composite clinical reference standard was used to establish case-level stone status; the source describes this composite standard but does not detail its exact components beyond that it served as the case-level ground truth.
Annotations for individual stones were reviewed by experts and used to evaluate localization completeness. The study design and data were retrospective; ethical oversight and approval were obtained (see Regulatory section).
Two deep-learning segmentation architectures were developed to delineate extrahepatic biliary anatomy and to detect and localize stones:
Both models were applied to representative IOC frames. The source does not provide additional training hyperparameters, augmentation strategies, or exact training dataset sizes beyond the held-out test set described below.
Model performance was evaluated on a held-out test set comprising 125 patients. Case-level diagnostic performance was measured against the composite clinical reference standard. Individual-stone localization was evaluated by comparing model-localized stones to expert-reviewed annotations.
Reported evaluation metrics at case level included sensitivity, specificity, and area under the receiver operating characteristic curve (AUC). Individual-stone localization was reported as the number of annotated stones successfully localized and the number of stone-positive cases in which all annotated stones were localized.
On the held-out test set of 125 patients, the segmentation models demonstrated the following case-level results against the composite clinical reference standard:
MiT-B2-UNet identified 23 of 25 stone-positive cases and 95 of 100 stone-negative cases. This corresponds to a sensitivity of 0.920, specificity of 0.950, and an AUC of 0.986.
nnU-Net identified 19 of 25 stone-positive cases and 98 of 100 stone-negative cases. This corresponds to a sensitivity of 0.760, specificity of 0.980, and an AUC of 0.959.
These results indicate both models achieved high specificity and overall discrimination (AUCs ≥ 0.959) with differing trade-offs between sensitivity and specificity: MiT-B2-UNet favored higher sensitivity, while nnU-Net favored higher specificity in this test set.
Individual-stone localization was less complete than case-level detection. Across annotated stones in the test set (59 annotated stones total), model localization counts were:
When considered per stone-positive case, models localized all annotated stones in a subset of cases:
These findings show that while models can often flag a case as stone-positive, complete localization of every annotated stone in a case was achieved in roughly half of stone-positive cases for both models.
The retrospective study was approved by the Cedars-Sinai Medical Center Institutional Review Board under protocol STUDY000001017, titled "AI-aided diagnosis of biliary anatomy," with a waiver of informed consent because the work used de-identified retrospective imaging and clinical data. The authors declared no competing interests.
Data supporting the findings are not publicly available due to patient privacy and institutional restrictions. The source reports that data may be available from the corresponding author on reasonable request and with appropriate institutional approvals. The article does not report public release of code or model weights; additional methodological details such as training hyperparameters and data partitioning beyond the held-out test set were not reported in the source.
In this retrospective, IRB-approved study, segmentation-based deep-learning models detected extrahepatic bile duct stones on representative IOC frames with high case-level discrimination (AUCs 0.959–0.986). MiT-B2-UNet showed higher sensitivity (0.920) while nnU-Net had higher specificity (0.980) in the reported test set. Individual-stone localization was more limited: models localized a portion of annotated stones (31/59 and 25/59) and fully localized all annotated stones in about half of stone-positive cases.
These results suggest that deep-learning segmentation approaches can assist in identifying stone-positive IOC cases and provide focal localizations of stones, but completeness of individual-stone localization varied and was not complete for many cases. The source does not provide prospective validation results or detailed deployment considerations. Data access is restricted but may be requested from the corresponding author with appropriate approvals.
Clinicians and imaging teams considering similar approaches should note the reported trade-offs between sensitivity and specificity across architectures and that individual-stone localization remains an area for further improvement and validation.