clustering taxonomyexploratory data analysispattern clusteringstatistical pattern recognitionunsupervised classification
Figures from the paper
Abstract (AI)
Clustering is the unsupervised classification of patterns (observations, data items, or feature vectors) into groups (clusters). The clustering problem has been addressed in many contexts and by researchers in many disciplines; this reflects its broad appeal and usefulness as one of the steps in exploratory data analysis. However, clustering is a difficult problem combinatorially, and differences in assumptions and contexts in different communities has made the transfer of useful generic concepts and methodologies slow to occur. This paper presents an overview of pattern clustering methods from a statistical pattern recognition perspective, with a goal of providing useful advice and references to fundamental concepts accessible to the broad community of clustering practitioners. We present a taxonomy of clustering techniques, and identify cross-cutting themes and recent advances. We also describe some important applications of clustering algorithms such as image segmentation, object recognition, and information retrieval.
Key Findings
1
Clustering is an unsupervised classification task that groups observations, data items, or feature vectors into clusters.
2
Clustering remains difficult because of combinatorial complexity and differing assumptions across research communities, which impede methodological transfer.
3
It develops a taxonomy of clustering techniques and identifies cross-cutting themes and recent advances.
4
The paper highlights clustering applications including image segmentation, object recognition, and information retrieval.
5
The paper provides a statistical pattern-recognition overview of clustering methods for a broad community of practitioners.
Research Object
Data clustering (unsupervised classification of patterns/observations/feature vectors into groups)
Research Subject
taxonomy, fundamental concepts, cross-cutting themes, recent advances, and applications of clustering techniques
Publication Details
Publication Date
1999-09-01
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
Artificial neural networks: a tutorial1996
Adaptation in Natural and Artificial Systems1992
Genetic algorithms in search, optimization, and machine learning1989
K-Means-Type Algorithms: A Generalized Convergence Theorem and Characterization of Local Optimality1984
Hierarchical Grouping to Optimize an Objective Function1963
Cited by12
The k-means Algorithm: A Comprehensive Survey and Performance Evaluation2020
dbscan: Fast Density-Based Clustering with R2019
A survey on Data Mining approaches for Healthcare2013
Standardization and Its Effects on K-Means Clustering Algorithm2013
Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts2013
Collinearity: a review of methods to deal with it and a simulation study evaluating their performance2012
Data mining: concepts and techniques2012
From Frequency to Meaning: Vector Space Models of Semantics2010
Understanding the concept of supply chain resilience2009
Segmentation of Multivariate Mixed Data via Lossy Data Coding and Compression2007
The Structure and Function of Complex Networks2003
An efficient k-means clustering algorithm: analysis and implementation2002