כתבה
arXiv cs.AI ·
Provable Benefit of SignGD: A Minimal Model Under Heavy-Tailed Class Imbalance
תקציר מקורי באנגליתarXiv:2512.00763v2 Announce Type: replace-cross Abstract: Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a large margin. Existing theory usually explains this gap by assuming favorable smoothness geometry or noise structure tailored to the specific optimizer. We instead ask whether such geometry can be induced from a concrete learning setting. Starting from an optimizer gap that persists across realistic language-modeling experiments, we progressively remove sequence dependence, architectural complexity, and stochasticity. We find that the gap exists in a minimal setting: the softmax unigram model with heavy-tailed data. This model exposes a simple deterministic mechanism under heavy-tailed class im
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית