An ensemble method for estimating the number of clusters in a big data set using multiple random samples

Abstract Clustering a big dataset without knowing the number of clusters presents a big challenge to many existing clustering algorithms. In this paper, we propose a Random Sample Partition-based Centers Ensemble (RSPCE) algorithm to identify the number of clusters in a big dataset. In this algorith...

Full description

Bibliographic Details
Main Authors: Mohammad Sultan Mahmud, Joshua Zhexue Huang, Rukhsana Ruby, Kaishun Wu
Format: Article
Language:English
Published: SpringerOpen 2023-04-01
Series:Journal of Big Data
Subjects:
Online Access:https://doi.org/10.1186/s40537-023-00709-4