International Journal of Biotechnology & Biochemistry
  • Year: 2008
  • Volume: 4
  • Issue: 2

A Novel Protein Coding Region Identifying Tool using Cellular Automata Classifier with Trust-Region Method and Parallel Scan Algorithm (NPCRITCACA)

  • Author:
  • P. Kiran Sree1,, I. Ramesh Babu2, N.S.S.N. Usha Devi
  • Total Page Count: 13
  • Page Number: 177 to 189

1Department of C.S.E, S.R.K.I.T, Vijayawada.

2C.S.E, Acharya Nagarjuna University, Guntur.

Abstract

Genes carry the instructions for making proteins that are found in a cell as a specific sequence of nucleotides that are found in DNA molecules. But, the regions of these genes that code for proteins may occupy only a small region of the sequence. Identifying the coding regions play a vital role in understanding these genes. A new measure for gene prediction in eukaryotes is presented. Analysis of all the experimental genes of S. cerevisiae revealed distribution of the phase in a bell-like curve around a central value, in all four nucleotides, whereas the distribution of the phase in the noncoding regions was found to be close to uniform. Similar findings were obtained for other organisms. Several measures based on the phase property are proposed. In protein coding regions, this rotation is assumed to closely align all vectors in the complex plane, thereby amplifying the magnitude of the vector sum. In noncoding regions, this operation does not significantly change this magnitude. Computing the measures with one chromosome and applying them on sequences of others reveals improved performance compared with other algorithms that use the 1/3 frequency feature, especially in short exons. The phase property is also used to find the reading frame of the sequence. In this paper we propose a NPCRITCACA tool for finding these protein coding regions which uses a Decision Tree Based classifier. This tool uses trust region method to find the closet (optimized) DNA nucleotide. It also uses a Sequential PRM-based protein folding algorithm for finding the point where these proteins add to the ladder. Cellular automata based parallel scan algorithm is used to provide parallel processing of the strides and its transitions. This proposed tool produces more accurate results, than that have previously been obtained for a range of different sequence lengths. Experimental results confirm the scalability of the proposed classifying tool to handle large volume of datasets irrespective of the number of classes, tuples and attributes. Good classification accuracy has been established.

Keywords

Decision Tree, Pattern Classifier, Sequential PRM-based protein folding algorithm, Standard parallel scan algorithm, Cellular Automata, trust region method