An evaluation of statistical approaches to text categorization

This paper is a comparative study of text categorization methods. Fourteen methods are investigated, based on previously published results and newly obtained results from additional experiments. Corpus biases in commonly used document collections are exami

An Evaluation of Statistical Approaches to Text Categorization yiming@cs.cmu.edu

Yiming Yang

April 10, 1997

CMU-CS-97-127

School of Computer Science Carnegie Mellon University Pittsburgh, PA 15213

Abstract This paper is a comparative study of text categorization methods. Fourteen methods are investigated, based on previously published results and newly obtained results from additional experiments. Corpus biases in commonly used document collections are examined using the performance of three classi ers. Problems in previously published experiments are analyzed, and the results of awed experiments are excluded from the cross-method evaluation. As a result, eleven out of the fourteen methods are remained. A k-nearest neighbor (kNN) classi er was chosen for the performance baseline on several collections; on each collection, the performance scores of other methods were normalized using the score of kNN. This provides a common basis for a global observation on methods whose results are only available on individual collections. Widrow-Ho, k-nearest neighbor, neural networks and the Linear Least Squares Fit mapping are the top-performing classi ers, while the Rocchio approaches had relatively poor results compared to the other learning methods. KNN is the only learning method that has scaled to the full domain of MEDLINE categories, showing a graceful behavior when the target space grows from the level of one hundred categories to a level of tens of thousands.

This research was supported in part by NIH grant LM-05714 and by NSF grant IRI9314992. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the o cial policies, either expressed or implied, of NIH or the U.S. Government.

An evaluation of statistical approaches to text categorization相关文档

最新文档

返回顶部