כתבה
arXiv cs.LG ·
GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization
תקציר מקורי באנגליתarXiv:2609.36074v1 Announce Type: new Abstract: Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for $K$-means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית