Noise-Robust Speaker Recognition System Using Deep Neural Architectures
Contributors
Arundhati Niwatkar
Sai Kiran Oruganti
Keywords
Proceeding
Track
Engineering, Sciences and Mathematics
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Speaker verification is recognized as a highly effective biometric authentication method thanks to its simplicity and high application range. Nowadays, speaker verification systems are widely implemented in voice-activated personal assistants, enhanced security systems, remote identity verification systems, and other devices for man-machine communication. Furthermore, a lot of innovations and techniques from deep learning have allowed to significantly increase verification accuracy by means of architectures such as x-vector embeddings, ECAPA-TDNN, and self-supervised speech representation models. Nevertheless, the performance of modern systems mostly suffers due to unfavorable conditions of actual operation, which include background noise, transmission channel inconsistencies, recording environment conditions, and speaker’s health problems such as illness or tiredness. Many researchers currently concentrate their efforts on improving the representation learning of the speakers by means of advanced neural networks, self-supervised learning methods and finding voice biomarkers.