Quantifying the Impact of Speaker and Content Features on ASR Systems Using Unsupervised Distance Metrics
Published in IEEE Sensors Reviews, 2025
Automatic speech recognition (ASR) models have become increasingly sophisticated, yet the underlying mechanisms driving their translation accuracy remain underexplored. This article explores the comparative influence of speaker characteristics and content similarity on ASR model performance, utilizing unsupervised distance metrics and clustering algorithms to gain deeper insights. By conducting a series of experiments using custom datasets, we aim to understand ASR model performance by examining whether the latent space features correlate more with speaker traits, such as accent, pitch, and speaking style, or with the semantic and syntactic content of the speech. Our findings reveal significant insights into the biases and strengths of current ASR technologies, highlighting the balance between speaker-dependent and content-dependent factors. Understanding these dynamics not only enhances the development of more robust and inclusive ASR systems but also paves the way for innovations in speech technology applications. This research contributes to the broader discourse on improving ASR models to better serve diverse populations and varied linguistic contexts.
Recommended citation: S. Pavuluri, S. De and A. K. Gupta, "Quantifying the Impact of Speaker and Content Features on ASR Systems Using Unsupervised Distance Metrics," IEEE Sensors Reviews, vol. 2, pp. 170-178, 2025
Download Paper
