יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אימון סטוכסטי סימטרי של שכבות התייחסות והבנה LoRA

Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA
במאמר זה, נוסחו ונוסחו תוצאות חדשות על אימון סטוכסטי של שכבות התייחסות והבנה LoRA. נוסחו גם תוצאות חדשות על קביעת קבועי Poincaré.
תקציר מקורי באנגליתarXiv:2605.07959v2 Announce Type: replace Abstract: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for a class of mild regularizations, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincar\'e inequality for the corresponding Gibbs' measure. Crucially, we show that the Poincar\'e constant is free of the data dim
קרא במקור המקורי