Published May 1, 2014 | Version v1

REDUCTION OF DIMENSION IN LINEAR BINARY CLASSIFICATION

  • 1. student of Saint-Petersburg state polytechnic university

Description

Consider the linear problem of binary classification (if the problem is linearly inseparable, it can be led to that by using a symmetric integral L-2 kernel). In solving this problem the classified elements are represented as elements of a vector space of dimension n. In practice n can be extremely large, for example for the task of classification of genes, it can reach tens of thousands. Large dimension implies high computation time and potentially high calculation error. Moreover, the use of large dimension can be expensive (for experimentation). So the question is: how to reduce n by discarding insignificant element components so that the elements were separated "no worse" in new space (were linearly separable) or "not much worse."

In this article I want to start with a brief overview of the technique from the article Isabel Guyon, Jason Weston, Stephen Barnhill, Vladimir Vapkin "Gene Selection for Cancer Classification using" and then offer my own method.

Notes

binary classification, reduction of dimension of data, the linear separability of sets, linear algebra, Farkas’s lemma, the system of linear inequalities, support vector machine, machine learning

Files

2_OsipyantsPM.pdf

Files (514.5 kB)

Name Size Download all
md5:c1d6c8b2a9acac364589c70e181f914f
514.5 kB Preview Download