כתבה
arXiv cs.LG ·
Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training
תקציר מקורי באנגליתarXiv:2505.15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic. Although this instability is often accompanied by attention-logit explosion, its underlying cause remains poorly understood. In this work, we present analytical insights and empirical evidence that dense local dependencies are a major contributor to attention-logit explosion. We demonstrate that dense local dependency patterns yield an effectively high-rank attention structure, which the low-rank parameterization of self-attention can only approximate with increasingly large logits as the sequence length grows. This logit inflation ultimately leads to training instability under low
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית