A Comparative Study of Acoustic Feature Extraction Using CSL and Web Prototype of LIS-N Application
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Purpose
Conventional clinical acoustic analysis systems are confined to specialized clinical environments, which limits scalable and remote voice monitoring. This study aimed to validate LIS-N, a web-based prototype of a novel mobile application that uses the open-source Praat-Parselmouth library for acoustic feature extraction, against the Computerized Speech Lab (CSL), the clinical reference standard, to assess its potential for telehealth and longitudinal voice assessment before integration into clinical or research workflows.
Method
Twenty adult volunteers without voice complaints or diagnosed dysphonia completed three voice tasks across three consecutive days at a tertiary academic voice center in quiet isolated rooms. All participants completed the full protocol across three consecutive days. Voice tasks included Rainbow Passage reading, sustained vowel /i/, and maximum phonation time /a/. Audio was captured concurrently using CSL and the LIS-N prototype, which uses Praat-Parselmouth for acoustic feature extraction. To isolate the contributions of recording hardware and analysis software, four cross-system conditions were also examined. Agreement was assessed using Pearson correlation and Bland-Altman analyses. Session-to-session trajectory correlations were computed to evaluate temporal consistency across systems.
Results
Fundamental frequency and Pitch Mean demonstrated excellent agreement across all tasks and conditions (r > 0.98), with robustness to hardware and software differences. Pitch Minimum and Pitch Maximum showed strong agreement during sustained vowel tasks. Energy-based features showed moderate agreement, with variability driven primarily by microphone differences, and same-microphone conditions yielded substantially stronger energy agreement. Perturbation measures including jitter, shimmer and CPP showed consistently poor cross-platform agreement, consistent with prior literature.
Conclusion
This study provides proof of concept for the viability of open-source acoustic frameworks as accessible alternatives to proprietary clinical systems, while also identifying limitations. Both systems should be interpreted separately, as each remains reliable within its own framework. The LIS-N prototype demonstrated strong agreement for frequency-based voice measures and consistent longitudinal monitoring, supporting adoption of open-source acoustic tools in telehealth and remote voice monitoring.