כתבה
arXiv cs.LG ·
How Bregman Divergences Shape Shampoo
תקציר מקורי באנגליתarXiv:2610.08534v1 Announce Type: new Abstract: Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences ca
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית