כתבה
arXiv cs.LG ·
NS-ATTENTION: טרנספורמציות ניוטון-שולץ של פלטי תשומת לב ב-Transformers חזותיים
NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers
NS-ATTENTION הוא אלגוריתם חדש שמשפר את דיוקן ה-Transformer החזותי. הוא מיישם טרנספורמציות ניוטון-שולץ על פלטי תשומת הלב, מה שמוביל לשיפור בדיוק. האלגוריתם נבדק על מודלים כמו ViT ו-Swin, והראה שיפורים משמעותיים בדיוק.
תקציר מקורי באנגליתarXiv:2609.27735v3 Announce Type: replace Abstract: Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a finite NS polynomial step and restore the original norm. The objective is to reduce spectral concentration and increase effective rank before standard head merging and output projection. Across ViT and Swin on CIFAR-10 and CIFAR-100, NS-Attn. impro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית