---
title: "OtoTymp-AI: Confidence-guided fusion of otoscopy video analysis and tympanometry for middle ear di"
id: "plos-one-5-confidence-guided-integration-of-otoscopy-videos-and-tympanometry-for-middle"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-5-confidence-guided-integration-of-otoscopy-videos-and-tympanometry-for-middle"
content_type: "clinical_feed_article"
specialty: "General"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357492"
published_at: "2026-09-01T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# OtoTymp-AI: Confidence-guided fusion of otoscopy video analysis and tympanometry for middle ear di
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-5-confidence-guided-integration-of-otoscopy-videos-and-tympanometry-for-middle
- **Specialty:** [General](https://medichelpline.com/clinical-feed/general.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357492)
- **Published At:** 2026-09-01T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- This multi-center study evaluated a confidence-guided multimodal framework, **OtoTymp-AI**, that integrates otoscopy video classification with conventional 226 Hz **tympanometry** to improve automated middle ear disease diagnosis. - The approach uses a convolutional neural network (**CNN**) to classify otoscopy videos and applies a clinically interpretable, rule-guided decision step: high-confidence video predictions are kept, while visually uncertain cases are refined using tympanometric rules. - Data came from three institutions: Nationwide Children’s Hospital for training, The Ohio State University for validation, and Vanderbilt University Medical Center as an external test set containing paired otoscopy videos and tympanometry. - Multimodal evaluation was performed on an external paired cohort of 104 videos across six diagnostic categories; the CNN-only model achieved 63.46% overall accuracy (95% CI, 53.88–72.08). - The hybrid video-plus-tympanometry framework reached a best-observed accuracy of 81.73% (95% CI, 73.22–87.98) at a confidence threshold of 0.90, an absolute improvement of 18.27 percentage points over the CNN alone (95% paired bootstrap CI, 9.62–26.92). - Noted improvements included specific diagnostic categories such as effusion and retraction, but per-class estimates had wide confidence intervals where sample sizes were small. - The study emphasizes the complementary value of anatomical video and physiological tympanometry data, and advocates for larger prospective paired multimodal cohorts to confirm optimal operating thresholds and per-class performance. - Source data and code are available via a public reproducibility package on GitHub; raw clinical videos and tympanometry records are restricted and available by request subject to IRB and data-use agreements.
## Clinical Analysis & Structured Key Points
Confidence-guided integration of otoscopy videos and tympanometry for middle ear disease diagnosis: A multi-center otoscopy study with external paired-cohort evaluation | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Reader Comments Figures Figures Abstract Accurate diagnosis of middle ear diseases using otoscopy remains challenging, particularly in primary care settings where clinician experience with otoscopy varies and visual examination alone may not fully capture middle ear physiology. Although AI-based models have shown promise for automated otoscopy interpretation, most rely primarily on visual information, creating an opportunity to improve diagnostic robustness by integrating complementary physiological measurements. Here, we present OtoTymp-AI, a confidence-guided multimodal decision-fusion framework that integrates otoscopy video analysis with conventional 226 Hz tympanometry. A convolutional neural network was trained to classify otoscopy videos, and tympanometric measurements were incorporated through a clinically interpretable, rule-guided decision strategy to refine visually uncertain predictions rather than through jointly trained multimodal representation learning. Within this multi-center otoscopy study, multimodal evaluation was performed in one independent external cohort comprising 104 videos with paired tympanometry across six diagnostic categories. In this exploratory paired-cohort analysis, the video-only CNN achieved an overall accuracy of 63.46% (95% CI, 53.88–72.08), while the hybrid video-plus-tympanometry framework achieved 81.73% accuracy (95% CI, 73.22–87.98) at the best observed threshold of 0.90, corresponding to an observed absolute improvement of 18.27 percentage points over the CNN-only model (95% paired bootstrap CI, 9.62–26.92). Improvements were observed for selected categories represented in this cohort, including effusion and retraction, although confidence intervals were wide for categories with small positive sample sizes. These findings suggest that integrating anatomical video information with physiological tympanometry may improve AI-assisted middle ear diagnosis in this external paired cohort, but the operating threshold and per-class performance require confirmation in larger prospective paired multimodal cohorts. Citation: Lu H, Langefeld CD, Zinnia A, Demir MF, Chavan S, Elmaraghy CA, et al. (2026) Confidence-guided integration of otoscopy videos and tympanometry for middle ear disease diagnosis: A multi-center otoscopy study with external paired-cohort evaluation. PLoS One 21(9): e0357492. https://doi.org/10.1371/journal.pone.0357492 Editor: Gauri Mankekar, LSU Health Shreveport, UNITED STATES OF AMERICA Received: May 13, 2026; Accepted: August 18, 2026; Published: September 1, 2026 Copyright: © 2026 Lu et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: The raw otoscopy videos, tympanometry records, and associated clinical metadata cannot be publicly shared because they are human-participant clinical data that may contain identifying or sensitive information even after de-identification, and public release is restricted by applicable institutional review board requirements and institutional data-use policies. Qualified researchers may submit requests for access to restricted de-identified data to Wake Forest University Health Sciences through its institutional Signing Official: Michele A. Gordon, CP, CRCP, Director of Research ( Michele.A.Gordon@AdvocateHealth.org ), Wake Forest University Health Sciences, Medical Center Boulevard, Winston-Salem, NC 27157-0001, USA. Requests will be considered subject to approval by the relevant participating institutions, applicable IRB requirements, and execution of any required data-use agreements. The public reproducibility package is available as: Lu H, Langefeld CD, Zinnia A, Demir MF, Chavan S, Elmaraghy CA, Wilson M, Niazi MKK, Moberly AC, Gurcan MN. otoscope_MMMC: code and anonymized derived data for “Confidence-guided integration of otoscopy videos and tympanometry for middle ear disease diagnosis: A multi-center otoscopy study with external paired-cohort evaluation.” GitHub; 2026. Available from: https://github.com/CAIR-LAB-WFUSM/otoscope_MMMC . Funding: This project was supported in part by the National Institute on Deafness and Other Communication Disorders of the National Institutes of Health under award number R01 DC020715 to Metin N. Gurcan and Aaron C. Moberly. The funder website is https://www.nidcd.nih.gov/ . There was no additional external funding received for this study. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health or the National Institute on Deafness and Other Communication Disorders. Competing interests: The authors have declared that no competing interests exist. Introduction Otoscopy is a fundamental, routinely performed examination for evaluating the external auditory canal, tympanic membrane, and middle ear. Ear-related complaints account for a substantial proportion of visits in primary care and pediatric settings, where clinicians are frequently required to assess ear pathology during both routine and acute care encounters [ 1 – 3 ]. However, access to otolaryngology specialists remains limited due to workforce shortages and long wait times, resulting in most initial evaluations being performed by non-specialist clinicians [ 2 , 4 , 5 ]. Despite its central role in the evaluation of ear disease, the diagnostic accuracy of clinicians performing otoscopy in primary care settings varies widely and is strongly influenced by clinician training and experience [ 6 , 7 ]. Visual interpretation of the tympanic membrane can be challenging, particularly in cases with subtle or overlapping features. Inaccurate diagnosis may lead to inappropriate treatment, delayed care, and adverse outcomes including hearing loss and reduced quality of life [ 8 ]. Recent advances in artificial intelligence (AI) have enabled the development of deep learning models for automated interpretation of otoscopic images and videos [ 9 – 12 ]. These models have demonstrated promising performance in controlled settings. However, most existing AI-based otoscopy models rely primarily on visual information and are often developed using single-center datasets [ 13 ], creating an opportunity to improve diagnostic robustness through external validation and integration of complementary clinical measurements. In particular, visual assessment alone may not fully distinguish conditions with similar appearance but different underlying middle-ear physiology. In clinical practice, otoscopy is often complemented by tympanometry [ 14 – 16 ], a non-invasive test that evaluates the mechanical properties of the tympanic membrane and middle ear system. Tympanometry provides objective physiological information that can help differentiate between conditions that are visually ambiguous, such as distinguishing effusion from a normal tympanic membrane. This complementarity creates an opportunity for AI-based diagnostic frameworks that integrate anatomical information from otoscopy with physiological information from tympanometry. In this study, we propose OtoTymp-AI, a confidence-guided multimodal decision-fusion framework that combines otoscopy video analysis with tympanometric measurements. Rather than learning a joint multimodal representation, OtoTymp-AI uses a clinically interpretable, confidence-triggered decision strategy: high-confidence video predictions are preserved, whereas visually uncertain predictions are refined using physiologically grounded tympanometry rules. This design was intended to evaluate whether complementary physiological information can improve AI-assisted middle ear diagnosis in a setting where paired otoscopy–tympanometry data are difficult to collect. To evaluate this framework, we conducted a retrospective multi-center study using otoscopy videos from three independent institutions. The Nationwide Children’s Hospital (NCH)-training dataset was used for model development, The Ohio State University (OSU)-validation dataset was used for model selection, and the Vanderbilt University Medical Center (VUMC)-external test dataset provided paired otoscopy videos and conventional 226 Hz tympanometry for multimodal evaluation. OtoTymp-AI was evaluated across six clinically relevant categories: effusion, normal, perforation, retraction, tympanostomy tube presence, and tympanosclerosis. This study demonstrates the potential of combining anatomical video information with physiological testing to improve AI-assisted middle ear diagnosis in an independent external clinical cohort. Accordingly, the multi-center design applies to otoscopy video model development, validation, and external testing, whereas the video-plus-tympanometry multimodal evaluation was performed in the single VUMC external paired cohort where ear-specific tympanometry was available. Related work AI-based interpretation of otoscopy images Early applications of AI in otology have primarily focused on static otoscopic images. Convolutional neural networks (CNNs) have been widely used to classify tympanic membrane conditions, with several studies [ 9 – 11 , 17 ] reporting high diagnostic accuracy under controlled conditions. These approaches typically rely on curated image datasets with well-centered views of the tympanic membrane and limited variability in acquisition settings. Despite promising results, static image–based methods have several practical limitations. Even a carefully selected single frame may not capture the entire tympanic membrane, particularly in pediatric otoscopy, where the narrower external auditory canal can restrict the field of view. As a result, diagnostically relevant findings may be visible only across multiple complementary views rather than in one image alone. In addition, many image datasets are collected from a single institution or under standardized acquisition conditions, raising concerns about generalizability to real-world clinical environments. Static images may also represent idealized views rather than the full variability encountered during routine otoscopy. Therefore, while static image models provide an important foundation, video-based analysis offers an opportunity to use multiple selected frames for more comprehensive diagnostic assessment. Video-based otoscopy analysis To address the limitations of static image analysis, more recent studies [ 18 – 20 ] have explored video-based approaches for otoscopic diagnosis. These methods either aggregate information across multiple frames or select diagnostically informative frames from video sequences to generate video-level predictions. By incorporating multiple viewpoints over time, video-based models can provide more complete coverage of the tympanic membrane—an advantage that is particularly important in pediatric otoscopy, where the restricted field of view often prevents a single image from capturing the entire eardrum. Video-based models can also better handle variability in image quality and reflect the dynamic nature of clinical examinations, making them more suitable for real-world deployment. However, several challenges remain for video-based otoscopy analysis. Existing approaches often focus on visual information alone, and performance is frequently evaluated using internal or randomly split datasets, with limited validation on independent external cohorts [ 21 ]. As a result, the robustness of video-based AI models across institutions, devices, and patient populations remains an important area for further study. In addition, visual information alone may not fully distinguish conditions with similar otoscopic appearance but different underlying middle-ear physiology. Multimodal approaches and tympanometry Tympanometry is a well-established clinical tool that provides objective assessment of tympanic membrane and middle ear function by measuring the mechanical properties of the tympanic membrane and middle ear system [ 22 ]. It is routinely used in conjunction with otoscopy to improve diagnostic accuracy, particularly in differentiating conditions such as middle ear effusion, perforation, and normal ear status [ 23 ]. Despite its clinical importance, tympanometry is not commonly incorporated into AI-based otoscopy frameworks. Prior work has explored decision fusion between image analysis and tympanometry [ 14 ], but multimodal AI approaches that combine otoscopy videos with tympanometric measurements remain limited. This represents an important opportunity because the two modalities provide complementary information: otoscopy captures anatomical appearance, whereas tympanometry reflects middle-ear mechanical function. Integrating these modalities may help refine visually uncertain predictions, particularly for conditions with overlapping otoscopic features but distinct physiological patterns. Together, these observations motivate an interpretable decision-fusion approach that uses tympanometry-derived physiological information to refine uncertain video-based AI predictions. This formulation targets transparent clinical decision support and is distinct from jointly trained multimodal representation learning. Materials and methods Data To evaluate model generalizability and prevent data leakage, the training, validation, and test datasets were strictly separated at both the institutional and patient levels. The NCH-training dataset was collected from Nationwide Children’s Hospital and used for model development. The OSU-validation dataset was collected from The Ohio State University and used for model selection. The VUMC-external test dataset was collected from Vanderbilt University Medical Center and included paired conventional 226 Hz tympanometry measurements collected as an additional research-related follow-up examination after the otoscopy video diagnosis had been completed. Tympanometry measurements were matched to the corresponding ear examination, enabling multimodal evaluation while preserving independence between the video-based reference diagnosis and tympanometry-derived model input. The cohort construction and institutional split are illustrated in Fig 1 , and the final analysis cohorts are summarized in Tables 1 and 2 . After quality control and high-quality frame selection, the final analysis cohort included 806 otoscopy videos from three institutions: 597 videos in the NCH-training dataset, 105 videos in the OSU-validation dataset, and 104 videos in the VUMC-external test dataset. Download: PNG larger image TIFF original image Fig 1. Dataset construction pipeline and cohort distribution across training, validation, and external test sets. Otoscopy videos were collected from three independent institutions and assigned to patient-disjoint cohorts. The NCH-training dataset underwent diagnostic agreement screening, video-level quality filtering, and automated high-quality frame selection for model development. The OSU-validation dataset was used for model selection without additional case-level exclusion, and the VUMC-external test dataset retained all otoscopy videos with available paired conventional 226 Hz tympanometry for multimodal evaluation. The final analysis cohort included 806 otoscopy videos: 597 training videos, 105 validation videos, and 104 external test videos. https://doi.org/10.1371/journal.pone.0357492.g001 Download: PNG larger image TIFF original image Table 1. Final dataset summary after quality control and high-quality frame selection. https://doi.org/10.1371/journal.pone.0357492.t001 Download: PNG larger image TIFF original image Table 2. Distribution of videos and selected frames by diagnostic category across the NCH-training, OSU-validation, and VUMC-external test datasets. https://doi.org/10.1371/journal.pone.0357492.t002 Reference standard labels were assigned as described in the “Reference Standard and Blinding” section below. Briefly, otoscopy videos were independently reviewed by at least two experienced clinicians, and the reference diagnosis was determined from otoscopic/video findings without use of tympanometry results. Cases with initial diagnostic disagreement were reviewed through consensus discussion; videos for which a consensus primary diagnosis could not be reached were excluded. Videos were also excluded if visualization of the tympanic membrane was severely compromised by motion artifacts, poor image quality, or obstruction by cerumen or debris, as determined through clinician review. The NCH-training dataset underwent training-cohort curation, including diagnostic agreement screening and video-level quality filtering, before automated high-quality frame selection. In contrast, the OSU-validation dataset was used without additional case-level exclusion, and the VUMC external paired cohort retained all otoscopy videos with available matched tympanometry. Automated high-quality frame selection was applied to videos for frame-level input preparation, but it did not remove additional OSU-validation or VUMC-external test cases from the analysis cohort. The primary analysis was performed at the ear-examination level. To maintain patient-disjoint dataset splits, all examinations from the same patient were assigned to only one dataset split; therefore, no patient contributed ear examinations to more than one split. Because some patients contributed bilateral ear examinations within the same split, we also summarized the number of unique patients and bilateral-ear contributors and performed patient-level sensitivity analyses to evaluate whether the main findings were sensitive to within-patient correlation. Fig 1 illustrates dataset curation and multi-institutional cohort composition. For the external test dataset, tympanometry measurements were paired with otoscopy videos after the clinician-assigned video reference diagnoses had been finalized. Matching between modalities was performed using study and clinical records to ensure correspondence to the same patient ear. Because tympanometry was collected as an additional research-related follow-up examination and required ear-specific matching, paired tympanometry was available only for a subset of the external cohort, contributing to the modest size of the paired multimodal test set. In the VUMC external paired cohort, the final analysis included 104 ear-level examinations from 76 unique patients. Twenty-eight patients contributed bilateral ear examinations, accounting for 56 ear-level examinations, whereas 48 patients contributed one ear examination. The cohort included 56 right-ear and 48 left-ear examinations. Although AOM cases were present in the NCH-training and OSU-validation datasets, no AOM cases were available in the VUMC external paired cohort after cohort curation. Therefore, external paired-cohort performance estimates were not calculated for AOM, and conclusions regarding multimodal performance should not be extended to AOM. The VUMC external paired cohort included 104 adult ear-level examinations. The median age was 51.0 years (IQR, 33.0–65.0; range, 18–85), and no pediatric cases were included. By age group, 43 examinations were from patients aged 18–44 years, 33 from patients aged 45–64 years, and 28 from patients aged ≥65 years. Reference standard and blinding The reference standard for model training and evaluation was defined at the ear-examination/video level using clinician-assig
## Related Clinical Research

- [Burden and adherence to long-term anti-VEGF intravitreal injections for retinal disease: qualitati](https://medichelpline.com/clinical-feed/bmj-open-11-treatment-burden-and-adherence-in-adults-receiving-long-term-anti-vegf.md)
- [Spatio-temporal imaging and acute manipulation of endogenous Yap1 in medaka using a Yap1-mGreenLan](https://medichelpline.com/clinical-feed/biorxiv-17-spatio-temporal-visualisation-and-manipulation-of-endogenous-yap1-in-medaka.md)
- [Non-retinoid RBP4 antagonist AKR-XI-85 as a candidate therapy for Stargardt disease](https://medichelpline.com/clinical-feed/biorxiv-19-a-non-retinoid-triazolopyrimidine-rbp4-antagonist-for-the-treatment-of.md)
- [Reduce Gun Access to Prevent Suicide: Evidence and Policy Options](https://medichelpline.com/clinical-feed/kff-health-news-0-it-s-hard-to-predict-who-will-be-suicidal-it-s-easier-to-ensure-people-can-t.md)
- [Kratom and 7‑OH dependence drives surge in addiction clinic visits](https://medichelpline.com/clinical-feed/stat-news-6-addiction-clinics-see-rising-cases-of-kratom-and-7-oh-withdrawal.md)

## Navigation
- [← Back to General Feed](https://medichelpline.com/clinical-feed/general.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.