The k-means Algorithm: A Comprehensive Survey and Performance Evaluation

Алгоритм k-means: всесторонний обзор и оценка производительности
Syed Mohammed Shamsul Islam, Mohiuddin Ahmed, Raihan Seraj
2020-08-12

experimental performance evaluationk-means clusteringk-means variantsnumber of clusters (k)random initialization
The k-means clustering algorithm is considered one of the most powerful and popular data mining algorithms in the research community. However, despite its popularity, the algorithm has certain limitations, including problems associated with random initialization of the centroids which leads to unexpected convergence. Additionally, such a clustering algorithm requires the number of clusters to be defined beforehand, which is responsible for different cluster shapes and outlier effects. A fundamental problem of the k-means algorithm is its inability to handle various data types. This paper provides a structured and synoptic overview of research conducted on the k-means algorithm to overcome such shortcomings. Variants of the k-means algorithms including their recent developments are discussed, where their effectiveness is investigated based on the experimental analysis of a variety of datasets. The detailed experimental analysis along with a thorough comparison among different k-means clustering algorithms differentiates our work compared to other existing survey papers. Furthermore, it outlines a clear and thorough understanding of the k-means algorithm along with its different research directions.
1
A fundamental limitation of k-means is its inability to handle various data types (i.e., data heterogeneity).
2
The paper surveys variants and recent developments of k-means that aim to overcome these shortcomings and discusses their effectiveness.
3
The work provides a detailed experimental analysis and comparative evaluation across multiple datasets, distinguishing it from prior survey papers.
4
k-means is a widely used and powerful clustering algorithm but has key limitations from random centroid initialization causing unstable convergence.
5
k-means requires the number of clusters to be specified in advance, which affects cluster shapes and sensitivity to outliers.

k-means clustering algorithm

Limitations, variants, recent developments, and empirical performance evaluation of the k-means algorithm (including initialization sensitivity, requirement to predefine number of clusters, handling of different data types, cluster shapes, and outlier effects)

Publication Details
Publication Date
2020-08-12
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Syed Mohammed Shamsul Islam
Mohiuddin Ahmed
Raihan Seraj
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%