כתבה
arXiv cs.AI ·
Misalignment of Low-Loss Regions Causes Grokking
תקציר מקורי באנגליתarXiv:2610.00620v1 Announce Type: cross Abstract: Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based on mode connectivity and the geometry of low-loss regions. The framework predicts that the standard modular-arithmetic setting does not always produce grokking: under a symmetry-preserving train/validation split, we observe a stable anti-grokking case in which validation performance does not recover. This counterexample challenges several existing correlational explanations of grokking. More broadly, our analysis framework and results further su
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית