Feature Selection via Correlation Coefficient Clustering

Hui-Huang Hsu, Cheng‐Wei Hsieh
2010-12-01

SCID:  54.1/zfkqcfss
Feature selection is a fundamental problem in machine learning and data mining. How to choose the most problem-related features from a set of collected features is essential. In this paper, a novel method using correlation coefficient clustering in removing similar/redundant features is proposed. The collected features are grouped into clusters by measuring their correlation coefficient values. The most class-dependent feature in each cluster is retained while others in the same cluster are removed. Thus, the most class-related and mutually unrelated features are identified. The proposed method was applied to two datasets: the disordered protein dataset and the Arrhythmia (ARR) dataset. The experimental results show that the method is superior to other feature selection methods in speed and/or accuracy. Detail discussions are given in the paper.
Publication Details
Publication Date
2010-12-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Hui-Huang Hsu
Cheng‐Wei Hsieh
Explore More Research
Use the citation graph to discover related papers and expand your research horizons.
Click any node to explore
Download PDF
100%