יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הגברת האיניציאליזציה של פעולות התנצבות

Conditioned Initialization for Attention
האיניציאליזציה של פעולות התנצבות במודלי Transformers עשויה להכניס שיפוט זמני באופטימיזציה. נציגים פתרון חדש לבעיה זו.
תקציר מקורי באנגליתarXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their success lies the attention layer, where the query, key, and value matrices determine how token dependencies are captured. While considerable work has focused on scaling and optimizing Transformers, comparatively little attention has been paid to how the weights of the queries, keys and values are initialized. Common practice relies on random initialization or alternatives such as mimetic initialization, which imitates weight patterns from converged models, and weight selection, which transfers weights from a teacher model. In this paper, we argue that initialization can introduce an optimization
קרא במקור המקורי