ZENITH International Journal of Multidisciplinary Research
  • Year: 2014
  • Volume: 4
  • Issue: 6

Studying preprocessing in character recognition system

  • Author:
  • Vikram Sharma
  • Total Page Count: 7
  • Page Number: 213 to 219

Asst. Professor, Dav College, Amritsar-143001, Punjab, India.

Online published on 13 August, 2014.

Abstract

An optical character recognition (OCR) is the process of converting scanned images of machine printed into computer readable codes. Recognition of characters of Indian languages like Hindi and Punjabi has been an area of research for many peoples and large numbers of research paper and reports have been published in this area. This generalized way for developing OCR system of any script involves preprocessing, segmentation, feature extraction & classification and Post processing. This paper describe mainly on preprocessing phase of OCR system. This stage takes digitized image (after scanning) as input, reduce noise and distortion, remove skew ness perform skeletonizing or thinning of the image and hence simplify the processing of remaining stages. An efficient noise removal algorithm is discussed in this paper. The fast and robust method for detecting and correcting skew angle is based on head line which is inherent properties of Indian languages like Hindi and Punjabi. For estimation of skew angle, the proper headline of the digitized image is detected and its skew angle is estimated with horizontal direction. An efficient thinning algorithm is discussed that mainly focus on the contour points of the given region. OCR system has never achieved a recognition rate that is 100% perfect and has certain limitation

Keywords

Preprocessing, Segmentation, Skeletonizing, Skew ness, Noise removal, Image loading