כתבה
arXiv cs.LG ·
אימון סטוכסטי סימטרי של שכבות התייחסות והבנה LoRA
Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA
במאמר זה, נוסחו ונוסחו תוצאות חדשות על אימון סטוכסטי של שכבות התייחסות והבנה LoRA. נוסחו גם תוצאות חדשות על קביעת קבועי Poincaré.
תקציר מקורי באנגליתarXiv:2605.07959v2 Announce Type: replace Abstract: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for a class of mild regularizations, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincar\'e inequality for the corresponding Gibbs' measure. Crucially, we show that the Poincar\'e constant is free of the data dim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית