Conventional clinical acoustic analysis systems are typically restricted to specialized environments, limiting scalable and remote voice monitoring. This study evaluated a web-based prototype, LIS-N, which uses the open-source Praat-Parselmouth library for acoustic feature extraction, against the clinical reference standard Computerized Speech Lab (CSL). The objective was to assess agreement across systems and the prototype's suitability for telehealth and longitudinal voice assessment prior to integration into clinical or research workflows.
Twenty adult volunteers without voice complaints or diagnosed dysphonia participated. Each completed three voice tasks across three consecutive days at a tertiary academic voice center. The voice tasks were:
All participants completed the full protocol across three consecutive days, and recordings were made in quiet isolated rooms.
Audio was captured concurrently using the clinical reference CSL and the web-based LIS-N prototype. LIS-N performs acoustic feature extraction via the open-source Praat-Parselmouth library. To isolate contributions from recording hardware versus analysis software, the study included four cross-system conditions, allowing comparison of same- and different-microphone and analysis-tool pairings.
Agreement between systems was assessed using Pearson correlation and Bland–Altman analyses. Session-to-session trajectory correlations were computed to evaluate temporal consistency across systems. Features evaluated included frequency-based measures (fundamental frequency and pitch statistics), energy-based measures, and perturbation measures (jitter, shimmer, and cepstral peak prominence [CPP]).
Fundamental frequency and Pitch Mean exhibited excellent agreement between CSL and LIS-N across all tasks and cross-system conditions, with Pearson correlation coefficients reported as greater than 0.98. These frequency-based metrics were robust to differences in both recording hardware and analysis software. Pitch Minimum and Pitch Maximum demonstrated strong agreement primarily during sustained vowel tasks, indicating that some pitch extrema may be task-dependent in their cross-system reliability.
Energy-based features achieved only moderate agreement overall. The study identified microphone differences as the primary driver of variability for energy measures; when the same microphone was used for both systems, agreement for energy-based metrics improved substantially. Perturbation measures — jitter, shimmer, and CPP — showed consistently poor cross-platform agreement, a finding consistent with prior literature and suggesting limited interchangeability for these measures between systems.
Session-to-session trajectory correlations were computed to evaluate temporal consistency. The LIS-N prototype demonstrated consistent longitudinal monitoring for frequency-based voice measures, supporting its potential for repeated remote assessments. The study reports that frequency metrics maintained reliable session-to-session trajectories across systems, reinforcing the feasibility of using open-source frameworks for longitudinal voice tracking.
This study provides proof of concept that an open-source acoustic framework implemented in a web-based prototype can approximate a clinical reference system for frequency-based measures of voice. The authors conclude that open-source tools like LIS-N using Praat-Parselmouth may be viable, accessible alternatives to proprietary clinical systems for telehealth and remote voice monitoring, particularly for measures such as fundamental frequency and Pitch Mean. However, the study also highlights limitations: energy-based measures are sensitive to microphone hardware, and perturbation measures (jitter, shimmer, CPP) do not show reliable cross-platform agreement. Accordingly, the authors recommend interpreting each system within its own framework rather than treating outputs as directly interchangeable.
The Institutional Review Board of the University of South Florida deemed this study exempt and classified it as not human research subject. The authors report that participant consent and applicable reporting guidelines were followed. Data supporting the findings are not publicly available due to participant privacy and confidentiality restrictions; access is limited to the approved study team and governed by institutional requirements. Declared funders include the NIH Common Fund (OT2-OD032720-01S3) and the University of South Florida.
The report emphasizes that while frequency-based measures showed excellent agreement, other feature classes did not, and that microphone and recording hardware exert substantial influence on energy measures. Perturbation metrics showed poor agreement consistent with previous work, indicating that these measures should be interpreted cautiously when comparing different systems. The study is a preprint and has not been peer reviewed; the authors note that its findings should not by themselves guide clinical practice. Further research and external validation would be required prior to clinical deployment or formal integration into care pathways.