כתבה
arXiv cs.LG ·
How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
תקציר מקורי באנגליתarXiv:2609.39767v1 Announce Type: new Abstract: The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית