Comparisons between dental artificial intelligence (AI) systems and dentists often involve analyzing how both interpret the same radiographic images. These paired comparisons are crucial for assessing whether AI can match or exceed human diagnostic accuracy. However, the methodology applied in many published studies tends to analyze AI and dentist results separately against a reference standard without considering their joint classification patterns. This overview explores the implications of this common practice and proposes improvements for more rigorous inference.
Most studies report diagnostic accuracy for dental AI and dentists independently but omit detailed joint classification information—specifically, how many cases both correctly or incorrectly diagnosed the same radiographs. Without this joint data, the exact distribution of concordant and discordant cases between AI and human readers remains unknown, limiting subsequent statistical analysis.
When the joint pattern of correct and incorrect classifications is missing, the difference in accuracy between AI and dentists can still be calculated; however, its uncertainty quantification, such as confidence intervals, cannot be reliably established. For example, in an analysis of 282 cases, the reported separate accuracies corresponded to 38 possible joint classification tables. These tables produced confidence intervals with widths varying by a factor of 2.5, highlighting substantial ambiguity in interpreting AI versus dentist performance.
The lack of precise joint data leads to a three-zone decision map framework when comparing accuracy differences:
This nuanced understanding aids more cautious interpretation of AI-dentist accuracy comparisons.
A key solution is simple: reporting the number of cases both AI and the dentist correctly classify. This single integer allows exact identification of the joint classification table, enabling the application of standard paired inference techniques. By including this measure, studies can produce more definitive assessments of whether AI truly surpasses or matches dentist performance.
In studies involving multiple readers, joint correct counts become even more critical. Pairwise dependencies among readers must come from a single joint distribution. This constraint means that joint correct counts involving a reference reader can tighten confidence bounds across all reader pairs. In one seven-reader dataset, introducing these joint counts reduced uncertainty bounds by a median of 37%, improving precision even for reader pairs that excluded the reference.
For scenarios where joint correct counts are not published, the authors developed DentalPair-Cert, a certification interval method. DentalPair-Cert provides finite-sample coverage guarantees uniformly over all admissible dependence structures between AI and dentist classifications under the independent sampling unit model. This approach certifies interval estimates through both nuisance maximization and inversion techniques, ensuring reliable inference despite uncertainty in joint behavior.
Extensive simulations exceeding 4.2 million comparisons demonstrated that ignoring paired dependence—treating AI and dentist results as independent—leads to coverage falling to 74.5% (below nominal levels) and type-I error inflation to 12.2%. These findings underscore the risk of misleading conclusions when joint classification data is omitted and paired dependencies unaccounted for.
A purposive sample of nine recent AI-dentist comparative studies revealed that only one reported any paired test based on discordant classification units. This highlights a pervasive gap in reporting standards and statistical rigor within the field of dental AI evaluation.
The study emphasizes the critical need for publishing joint correct classification counts alongside separate diagnostic accuracies in dental AI versus dentist studies. Doing so enables rigorous paired statistical inference to accurately assess AI performance. The DentalPair-Cert methodology offers a certified approach to inference when such joint data is unavailable, improving reliability and reproducibility in this emerging research area. Enhanced reporting standards adopting these recommendations will strengthen the clinical evaluation of dental AI systems moving forward.