כתבה
arXiv cs.AI ·
גרדיאנטים משוקללים חשפים את חשיבות הפרמטרים ואת סוגי השגיאות ב-LLMs
Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs
גרדיאנטים משוקללים חשפים את חשיבות הפרמטרים ואת סוגי השגיאות ב-LLMs. ניתן לראות כיצד גרדיאנטים משוקללים יכולים לסייע בהבנת חשיבות הפרמטרים ובזיהוי סוגי השגיאות ב-LLMs.
תקציר מקורי באנגליתarXiv:2607.10803v2 Announce Type: replace-cross Abstract: Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter importance that explicitly captures the interaction between model weights and first-order gradient information and identifies parameters that disproportionately influence model behavior, such as those responsible for collapse phenomena in LLMs. Across a range of models and settings, we show that WAG surfaces a tiny but critical subset of parameters (< 0.5 parts per million or 0.00005% of model size) whose modification leads to dramatic degradation in performance, indicating a no
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית