Voice Activity Detection Based on SNR and Non-Intrusive Speech Intelligibility Estimation

Research output: Contribution to journalArticlepeer-review

Abstract

This paper proposes a new voice activity detection (VAD) method which is based on SNR and non-intrusive speech intelligibility estimation. In the conventional SNR-based VAD methods, voice activity probability is obtained by estimating frame-wise SNR at each spectral component. However these methods lack performance in various noisy environments. We devise a hybrid VAD method that uses non-intrusive speech intelligibility estimation as well as SNR estimation, where the speech intelligibility score is estimated based on deep neural network. In order to train model parameters of deep neural network, we use MFCC vector and the intrusive speech intelligibility score, STOI (Short-Time Objective Intelligent Measure), as input and output, respectively. We developed speech presence measure to classify each noisy frame as voice or non-voice by calculating the weighted average of the estimated STOI value and the conventional SNR-based VAD value at each frame. Experimental results show that the proposed method has better performance than the conventional VAD method in various noisy environments, especially when the SNR is very low.
Original languageEnglish
Pages (from-to)26-30
Number of pages5
JournalThe International Journal of Internet, Broadcasting and Communication
Volume11
Issue number4
DOIs
StatePublished - Nov 2019

Fingerprint

Dive into the research topics of 'Voice Activity Detection Based on SNR and Non-Intrusive Speech Intelligibility Estimation'. Together they form a unique fingerprint.

Cite this