*Professor & Head, Department of EIE, SNS College of Technology, Coimbatore, India
**Professor, Department of Mechanical Engineering, SNS College of Technology, Coimbatore, India
***Assistant Professor, SNS College of Technology, Coimbatore, India
Online published on 1 June, 2016.
Speech To Text (STT) conversion systems are widely used in Automatic Voice Response Systems and speech to text converters. Accurate STT requires efficient feature extraction and faster classification methodologies. This work presents MFCC based feature extraction for language model and adaptive ensemble ELM for classification of trained word. A set of ten English words were recorded by a single speaker with each word pronounced and recorded 20 times. The recorded voice templates of each word were preprocessed using a denoising filter. Upon denoising the features of speech were extracted by using MFCC. An ELM based training was adopted for making a language model and classification. ELM based classification on such a small data set gave 79.5% of accuracy with faster classification time. The work can be further extended by extending the database and testing for a speaker independent system.
MFCC, ELM, PIEX, Speech to Text (STT), Filter