International Journal of Managment, IT and Engineering
  • Year: 2013
  • Volume: 3
  • Issue: 6

ScanDroid: Automatic classification of document images on android mobile devices

  • Author:
  • Gunjan Joshi, Prateek Bedmutha, Aman Patel, Kinjal Bathani
  • Total Page Count: 10
  • Page Number: 528 to 537

Department of Computer Engineering, Sinhgad College of Engineering, University of Pune, India

Online published on 7 November, 2013.

Abstract

Text categorization refers to the automatic labelling of documents, based on natural language text contained in or associated with each document, into one or more pre-defined categories. Today, image categorization is a necessity due to a very large amount of image documents that we have to deal with daily. The current image categorization system uses an associated text approach for classification of images. We propose herein a new approach for automatic image categorization on android mobile devices, an application for classification of document images based on its contents, which is useful to businessmen, teachers and students. The classification module is the primary module. OCR technology is used to extract the textual contents from the input images. The textual contents extracted are given as input to the classification module which automatically classifies the images based on hashing techniques. The searching module is used to search for relevant image documents based on user keyword. The interface of the Android OS makes the end-user easy and efficient to search the relevant images into the database based on user keyword. Our literature survey leads to conclusion that mining is a good and promising strategy for automatic image categorization.

Keywords

Text Mining, Information Retrieval, Data Mining, OCR, Text classification, Feature extraction