Noisy speech emotion recognition using sample reconstruction and multiple-kernel learning

Article ID	Journal	Published Year	Pages	File Type
7116809	The Journal of China Universities of Posts and Telecommunications	2017	10 Pages	PDF

Abstract

Speech emotion recognition (SER) in noisy environment is a vital issue in artificial intelligence (AI). In this paper, the reconstruction of speech samples removes the added noise. Acoustic features extracted from the reconstructed samples are selected to build an optimal feature subset with better emotional recognizability. A multiple-kernel (MK) support vector machine (SVM) classifier solved by semi-definite programming (SDP) is adopted in SER procedure. The proposed method in this paper is demonstrated on Berlin Database of Emotional Speech. Recognition accuracies of the original, noisy, and reconstructed samples classified by both single-kernel (SK) and MK classifiers are compared and analyzed. The experimental results show that the proposed method is effective and robust when noise exists.

Keywords

Multiple-kernel learning Feature selection Compressed sensing Speech emotion recognition