AI‑Driven Acoustic Feature Extraction for Automated Stuttering Detection
Contributors
Dr. Salma Jabeen
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Stuttering is recognized as a speech fluency disorder characterized by involuntary repetitions, prolongations, and blocks occurring during natural speech production [1]. Conventional clinical diagnosis predominantly depends on perceptual assessment by speech‑language professionals, which may introduce subjectivity and inter‑evaluator variability [1]. With significant progress in speech signal processing and the emergence of artificial intelligence (AI), automated and objective approaches to stuttering detection have become increasingly viable [2], [3].This paper provides an overview of AI‑based approaches for extracting acoustic features used in automated stuttering detection, with particular emphasis on Mel‑Frequency Cepstral Coefficients (MFCCs) and deep learning models [2], [4]. A comparative analysis of existing literature indicates that deep neural networks (DNNs), including CNNs, LSTMs, and transformer‑based architectures, consistently outperform traditional machine learning algorithms when applied to disfluent speech data [3], [5]–[8]. Despite these promising outcomes, several limitations persist, primarily due to speaker variability, imbalanced datasets, and the high computational complexity associated with deep learning‑based approaches [6], [7], [9]–[11].