A Novel Self-Supervised Transformer Model for Speech Disfluency Boundary Detection and Waveform Reconstruction


Date Published : 4 August 2026

Contributors

Shashi Kant Gupta

Lincoln University College, Malaysia
Author

Keywords

Speech Remediation Wav2Vec 2.0 Surgical Excision SEP-28k Disfluency Detection Concatenative Synthesis Assistive Technology

Proceeding

Track

Engineering, Sciences and Mathematics

License

Copyright (c) 2026 Sustainable Global Societies Initiative

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Abstract

Speech technology is getting better at finding more than solving the problem of stuttering. Disfluency is actually fixing by remediation. The main goal of the paper is a Wav2Vec2 2.0 model fine-tuned on SEP-28k dataset. in which we use a probability threshold of 0.93 to locate boundaries in disfluency like block prolongation, repetition within a 16KHz audio stream. Once the flagged segment is caught it gets cut and the audio is stitched using a 30ms Hamming window crossfade. We run evaluation across noise conditions between 5dB to 25dB SNR. Spectral distortion at signal level is tracked by LPC analysis. For Linguistic Integrity Whisper-Base is used as a secondary auditor with Word Error Rate target below 0.15. Accuracy is 92% of detection and 99% of recall in the listening study of 40 participants. The mean of quality rating is ranging from 2.2 to 4.3 vocal identity and semantics of content remained intact.

References

No References

Downloads

How to Cite

Gupta, P. (Dr.) S. K. G. (2026). A Novel Self-Supervised Transformer Model for Speech Disfluency Boundary Detection and Waveform Reconstruction. Sustainable Global Societies Initiative, 1(10). https://vectmag.com/sgsi/paper/view/892