面向中爪哇省大数据贫困分析的混合分裂式K均值框架

Authors

  • 朱思雨 (通讯作者) 贵州大学经济学院,贵阳 550000,贵州,中国

关键词:

大数据; 聚类; 分裂式层次聚类; 混合模型; K均值; 贫困数据分析

摘要

聚类在大数据分析中不可或缺,对于划分高维社会经济数据集以支撑结果解读与政策决策尤为重要。K均值因简单且可扩展而被广泛使用,但它对初始质心选择高度敏感,常导致结果不稳定与收敛缓慢。以往的混合方法(如凝聚式—K均值)试图用层次聚类完成质心初始化来解决该问题,然而这类方法依赖自下而上的合并,可能产生次优的初始划分,并在数据规模增大时增加计算开销。为克服上述局限,本文提出一种混合分裂式—K均值(DHC)模型:先以自上而下的层次分裂生成更为一致的初始质心,再由K均值加以精化。本文使用印度尼西亚中央统计局(BPS)提供的中爪哇省多维贫困数据集,把DHC与标准K均值、凝聚式—K均值进行了对比评价,评价内容包括执行时间、收敛迭代次数与聚类有效性指标(轮廓系数、Davies–Bouldin指数与Calinski–Harabasz指数)。实验结果表明,DHC使执行时间最多降低97%,所需迭代次数比标准K均值少40%,同时聚类质量相当甚至更优(如CH指数由14.3提高到15.8)。上述结果表明,DHC模型为大规模社会经济数据分析提供了一种更高效、更稳定的聚类方案。

Abstract

Clustering is essential in big data analytics, especially for partitioning high-dimensional socioeconomic datasets to support interpretation and policy decisions. While K-Means is widely used for its simplicity and scalability, its strong sensitivity to initial centroid selection often leads to unstable results and slower convergence. Previous hybrid approaches, such as Agglomerative-K-Means, attempted to address this issue by using hierarchical clustering for centroid initialization; however, these methods rely on bottom-up merging, which can produce suboptimal initial partitions and increase computational overhead for larger datasets. To overcome these limitations, this study proposes a hybrid divisive-K-Means (DHC) model that employs top-down hierarchical splitting to generate more coherent initial centroids before refinement with K-Means. Using a multidimensional poverty dataset from Central Java Province provided by the Indonesian Central Bureau of Statistics (BPS), the performance of DHC was evaluated against standard K-Means and Agglomerative-K-Means. The assessment included execution time, convergence iterations, and cluster validity indices (Silhouette, Davies-Bouldin, and Calinski-Harabasz). Experimental results demonstrate that DHC reduces execution time by up to 97% and requires 40% fewer iterations than standard K-Means, while achieving comparable or improved cluster quality (e.g., CH Index increasing from 14.3 to 15.8). These findings indicate that the DHC model offers a more efficient and stable clustering solution for large-scale socioeconomic data analysis.

References

[1] Ezugwu A E, Elsisi S, Hussain A G, et al. Hybrid firefly algorithms for clustering. IEEE Access, 2020, 8: 121089-121118.

[2] Ahmed M, Seraj R, Islam S M S. The k-means algorithm: a comprehensive survey and performance evaluation. Electronics, 2020, 9(8): 1295.

[3] Ezugwu A E, Akinola I A, Elsisi M H, et al. A comprehensive survey of clustering algorithms. Engineering Applications of Artificial Intelligence, 2022, 110: 104743.

[4] 武福林. 深度聚类算法研究: 方法、评估与应用. 信息记录材料, 2026(6).

[5] 赵杰, 雷秀娟, 吴振强. 基于最优类中心扰动的萤火虫聚类算法. 计算机工程与科学, 2015, 37(2): 314-320.

[6] Bai L, Liang J, Cao F. A multiple k-means clustering ensemble algorithm to find nonlinearly separable clusters. Information Fusion, 2020, 61: 36-47.

[7] Chen Y, Wang L, Li J. A clustering-based approach using K-means and hierarchical methods for large-scale data. Information, 2025, 16(6): 441.

[8] Zhang Y. A comprehensive survey on traffic missing data imputation. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(2): 1234-1245.

[9] 贺玲, 吴玲达, 蔡益朝. 数据挖掘中的聚类算法综述. 计算机应用研究, 2007(1): 10-13.

Downloads

发布日期

2026-08-08

How to Cite

朱思雨. 面向中爪哇省大数据贫困分析的混合分裂式K均值框架. 现代工程与应用. 2026, 4(3): 1-8. DOI: https://doi.org/10.61784/mea2026.