Journal of Wavelet Theory and Applications
  • Year: 2008
  • Volume: 2
  • Issue: 1

Speech Recognition using Wavelet Packet Features

  • Author:
  • Mihalis Siafarikas, Iosif Mporas, Todor Ganchev, Nikos Fakotakis
  • Total Page Count: 19
  • Page Number: 41 to 59

Wire Communications Laboratory, Department of Electrical and Computer Engineering, University of Patras, 26500 Rion-Patras, Greece.

Abstract

In view of the growing use of automatic speech recognition in the modern society, we study various alternative representations of the speech signal that have the potential to contribute to the improvement of the recognition performance. Specifically, the main targets of the present article are to overview and evaluate the practical importance of some recently proposed, and thus less studied, wavelet packet-based speech parameterization methods on the speech recognition task, illustrating their merits compared to other well known approaches. To this end, working on the widely acknowledged TIMIT (Texas Instruments and Massachusetts Institute of Technology) speech database and relying on the Sphinx-III speech recognizer, we contrast the performance of four wavelet packet-based speech parameterizations against traditional Fourier-based techniques that have been considered for the task of speech recognition for over two decades, including Mel Frequency Cepstral Coefficients (MFCC) and Perceptual Linear Predictive (PLP) cepstral coefficients that presently dominate the speech recognition field. The experimental results demonstrate that the wavelet packet-based speech features of interest provide a superior performance over the baseline parameters. This validates the wavelet packet-based speech parameterization schemes as a promising research direction that could bring further reduction of the speech recognition error rate.

Keywords

speech parameterization, speech recognition, wavelet packet transform