An Investigation into Using Parallel Data for Far-Field Speech Recognition

Yanmin Qian; Tian Tan; Dong Yu

An Investigation into Using Parallel Data for Far-Field Speech Recognition

Yanmin Qian ,
Tian Tan ,
Dong Yu

April 2016

Published by IEEE - Institute of Electrical and Electronics Engineers

Download BibTex

Far-ﬁeldspeechrecognitionisanimportantyetchallengingtaskdue to low signal to noise ratio. In this paper, three novel deep neural network architectures are explored to improve the far-ﬁeld speech recognition accuracy by exploiting the parallel far-ﬁeld and closetalk recordings. All three novel architectures use multi-task learning for the model optimization but focus on three different ideas: dereverberation and recognition joint-learning, close-talk and farﬁeld model knowledge sharing, and environment-code aware training. Experiments on the AMI single distant microphone (SDM) task show that each of the proposed method can boost accuracy individually, and additional improvement can be obtained with appropriate integration of these models. Overall we reduced the error rate by 10% relatively on the SDM set by exploiting the IHM data.

© IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other users, including reprinting/ republishing this material for advertising or promotional purposes, creating new collective works for resale or redistribution to servers or lists, or reuse of any copyrighted components of this work in other works.