Published April 30, 2023 | Version CC BY-NC-ND 4.0
Journal article Open

Hybrid Approach to Detect Prolonged Speech Segments

  • 1. Research Scholar, Visvesvaraya Technological University (VTU), Belgaum (Karnataka), India.
  • 2. Associate Professor, Department of Computer Science and Engineering, JSS Science and Technology University, Mysore (Karnataka), India.

Contributors

Contact person:

  • 1. Research Scholar, Visvesvaraya Technological University (VTU), Belgaum (Karnataka), India.

Description

Abstract: In the last 10 decades various methods have been introduced to detect prolonged speech segments automatically for stuttered speech signals. However less attention has been paid by researches in the detection of prolongation disorder at the parametric level. The aim of this study is to propose a hybrid approach to detect the prolonged speech segments by combining various spectral parameters with their recognition accuracies for the reconstructed speech signal. The paper presents prolonged segments detection by considering the parameters individually, combining various spectral parameters, validation of prolongation detection system, MFCC feature extraction process, basic model accuracies for the reconstructed signals. The proposed methods are simulated and experimented on UCLASS derived dataset. Obtained results are compared with the existing works of prolongation detection at parametric and word level. It is observed that hybrid parameters yield 92% of recognition rate for larger frame sizes of 200ms when modeled with SVM. The results are also tabulated and discussed for various metrics like sensitivity, specificity and accuracy metrics in detecting the prolonged segments. The study also focuses on the prolongation characteristics of vocalized and non-vocalized sounds at phoneme level. The detection accuracy of 71% is observed for Vocalized prolonged vowel phonemes over non-vocalized prolonged signal. Objectives: The objective of this work is to propose a hybrid algorithm to detect prolonged segments automatically for speech signal with prolongation disorder. The other objective is to evaluate the obtained spectral parameters performances by applying to various evaluation metrics and models to compute the recognition accuracy of a reconstructed signal. The objective is further extended to bring out the importance of variable frame size concept and to analyze the variations in vocalized and non-vocalized sounds. Methods: The methods adopted to detect prolonged speech segments are discussed at two levels namely at the preprocessing and modeling levels. The Preprocessing level is discussed by applying various parameters at an individual level, hybrid level by combing the Centroid, Entropy, Energy, ZCR parameters and MFCC feature extraction method. A new method has been applied using Specificity, Sensitivity and accuracy metrics to validate the prolongation detection model performance. In modeling level, the above parameters are discussed by applying evaluation metrics for the clustering and classification models like K-means, FCM and SVM. The performance of these methods is considered for evaluating and estimating the prolonged segment detection accuracy of the reconstructed speech signals of vocalized and non-vocalized sounds. All these methods are discussed in detail in the following sections. Findings: Hybridizing the spectral parameters to detect the prolonged speech segment automatically is a major finding of this work. It is also found that Specificity, sensitivity and accuracy metrics plays a major role in designing and validating the prolongation detection model. From the further experiments it is identified that the hybrid and verification metrics suits better for vocalized and non-vocalized sounds when larger frame lengths are considered. SVM has been found to perform better for all the above considerations. Novelty: As per Literature survey it is observed that individual and few parameters are applied to detect the prolongation. But works are not addressed on applying or combining more than two parameters to detect the prolonged speech segments. The novelty of this work lies in selecting and combining the spectral parameters at the preprocessing stage to detect the prolongation disorder. Spectral centroid and entropy are considered as appropriate parameters along with ZCR and Energy parameters. Hence hybridizing these parameters results in a novelty to propose an automatic prolongation detection system. Novelty is further brought by applying Specificity, sensitivity and accuracy metrics to build and evaluate the detection system for vocalized and non-vocalized prolonged sounds.

Notes

Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved.

Files

D41060412423.pdf

Files (675.8 kB)

Name Size Download all
md5:e3d6cb554c1df7d9675494e12a574aa7
675.8 kB Preview Download

Additional details

Related works

Is cited by
Journal article: 2249-8958 (ISSN)

References

  • Dr.MA Anusuya, SK Katti, Front end analysis of speech recognition: a review", International Journal of Speech Technology,2011, Available from: DOI:10.1007/s10772-010-9088-7
  • K B Drakshayini, Anusuya M A," Stop gap removal using spectral parameters for stuttered speech signal", International Journal of Advanced Trends in Computer Science and Engineering ,2021 Available fromhttps://doi.org/10.30534/ijatcse/2021/521032021
  • Om Dadaji Deshmukh, Suraj Satishkumar Sheth, Ashish Verma," Reconstruction of a smooth speech signal from a stuttered Speech",2013 Available from: https://patentimages.storage.googleapis.com/69/64/27/5aef3d5d69024c/US8600758.pdf
  • Katarzyna Barczewska, Magdalena Igras-Cybulska, "Detection of disfluencies in speech signal",2013 Available from: https://www.researchgate.net/publication/261913703_Detection_of_disfluencies_in_speech_signal
  • G. Manjula, M. Shiva Kumar, "Identification and Validation of Repetitions/Prolongations in Stuttering Speech using Epoch Features",2017 Available from: https://www.ripublication.com/ijaer17/ijaerv12n22_29.pdf
  • Sadeen Alharbi, Madina Hasan, Anthony J H Simons, Shelagh Brumfitt , Phil Green,"A Lightly Supervised Approach to Detect Stuttering in Children's Speech",2018, Available from: https://www.isca-speech.org/archive/pdfs/interspeech_2018/alharbi18_interspeech.pdf Available from: DOI:10.1016/j.procs.2020.04.146
  • Waldemar Suszyńskia, Wiesława Kuniszyk-Jóźkowiaka, Elżbieta Smołkaa, Mariusz Dzieńkowskib "Prolongation detection with application of fuzzy logic",2021, Available from: https://core.ac.uk/download/235272168.pdf
  • Barlian Henryranu Prasetio, Edita Rosana Widasari, Hiroki Tamura," Multiscale-based Peak Detection on Short Time Energy and Spectral Centroid Feature Extraction for Conversational Speech Segmentation",2021, ICPS Proceedings, SIET '21 Available from: https://dl.acm.org/doi/10.1145/3479645.3479675
  • Sakshi Gupta, Ravi S. Shukla, Rajesh K. Shukla, Rajesh Verma," Deep Learning Bidirectional LSTM based Detection of Prolongation and Repetition in Stuttered Speech using Weighted MFCC",2020, IJACSA, available form:10.14569/IJACSA.2020.0110941
  • Vinay N A, Bharathi S H," Dysfluency Recognition by using Spectral Entropy Features", IJEAT,2019 Available from: https://www.ijeat.org/wp-content/uploads/papers/v8i6/F7881088619.pdf
  • Manjutha M, Subashini P," Statistical Model-Based Tamil Stuttered Speech Segmentation Using Voice Activity Detection",2022, Journal of Positive School Psychology, Available form https://journalppw.com/index.php/jpsp/article/view/9958/6481
  • Salsabil Besbes, Zied Lachiri," Multi-class SVM for stressed speech recognition",2016, IEEE conference proceedings, Available from: DOI: 10.1109/ATSIP.2016.7523188
  • H.Y. Vani, Dr.M.A. Anusuya and Dr.M.L. Chayadevi "Fuzzy Clustering Algorithms - Comparative Studies for Noisy speech signals", 2019, ICTACT Journal on Soft Computing, Available from: https://ictactjournals.in/paper/IJSC_Vol_9_Iss_3_Paper_5_1920_1926.pdf
  • P. Howell, S. Davis, and J. Bartrip, "The university college London archive of stuttered speech (Uclass)", Journal of Speech, Language, and Hearing Research, vol. 52, pp. 556–569, 2009. Available from: DOI: 10.1044/1092-4388(07-0129)

Subjects

ISSN: 2249-8958 (Online)
https://portal.issn.org/resource/ISSN/2249-8958#
Retrieval Number: 100.1/ijeat.D41060412423
https://www.ijeat.org/portfolio-item/D41060412423/
Journal Website: www.ijeat.org
https://www.ijeat.org
Publisher: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP)
https://www.blueeyesintelligence.org