A Novel Self-Supervised Transformer Model for Speech Disfluency Boundary Detection and Waveform Reconstruction
Contributors
Shashi Kant Gupta
Keywords
Proceeding
Track
Engineering, Sciences and Mathematics
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Speech technology is getting better at finding more than solving the problem of stuttering. Disfluency is actually fixing by remediation. The main goal of the paper is a Wav2Vec2 2.0 model fine-tuned on SEP-28k dataset. in which we use a probability threshold of 0.93 to locate boundaries in disfluency like block prolongation, repetition within a 16KHz audio stream. Once the flagged segment is caught it gets cut and the audio is stitched using a 30ms Hamming window crossfade. We run evaluation across noise conditions between 5dB to 25dB SNR. Spectral distortion at signal level is tracked by LPC analysis. For Linguistic Integrity Whisper-Base is used as a secondary auditor with Word Error Rate target below 0.15. Accuracy is 92% of detection and 99% of recall in the listening study of 40 participants. The mean of quality rating is ranging from 2.2 to 4.3 vocal identity and semantics of content remained intact.