Published November 7, 2018 | Version v1

Dimensionality Reduction and Feature Selection Methods for Script Identification on Document Images

Description

Abstract

The goal of this research is to explore effects of dimensionality reduction and feature selection on the problem of script identification from images of printed documents. The kadjacent segment is ideal for this use due to its ability to capture visual patterns. We have used principle component analysis to reduce the size of our feature matrix to a handier size that can be trained easily, and experimented by including varying combinations of dimensions of the super feature set. A modular approach in neural network was used to classify 7 languages – Arabic, Chinese, English, Japanese, Tamil, Thai and Korean.

Keywords

Feature reduction; feature selection; neural networks; principle component analysis; script identification

For More Details: http://it-in-industry.com/itii_papers/2014/2114itii01.pdf

Files

2114itii01.pdf

Files (296.5 kB)

Name Size Download all
md5:543d57c16d717feef9465b8ba304d395
296.5 kB Preview Download