Article ID: | iaor20119608 |
Volume: | 62 |
Issue: | 7 |
Start Page Number: | 2793 |
End Page Number: | 2800 |
Publication Date: | Oct 2011 |
Journal: | Computers and Mathematics with Applications |
Authors: | Meng Jiana, Lin Hongfei, Yu Yuhai |
Keywords: | datamining |
Feature selection for text categorization is a well‐studied problem and its goal is to improve the effectiveness of categorization, or the efficiency of computation, or both. The system of text categorization based on traditional term‐matching is used to represent the vector space model as a document; however, it needs a high dimensional space to represent the document, and does not take into account the semantic relationship between terms, which leads to a poor categorization accuracy. The latent semantic indexing method can overcome this problem by using statistically derived conceptual indices to replace the individual terms. With the purpose of improving the accuracy and efficiency of categorization, in this paper we propose a two‐stage feature selection method. Firstly, we apply a novel feature selection method to reduce the dimension of terms; and then we construct a new semantic space, between terms, based on the latent semantic indexing method. Through some applications involving the spam database categorization, we find that our two‐stage feature selection method performs better.