כתבה
arXiv cs.LG ·
למידה במצב חוזר: גרדיאנט דסנט עם רשתות רצורתיות לינאריות
Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks
במאמר זה, המחברים מציגים רשת רצורתית לינארית שמסוגלת ללמוד במצב חוזר דרך גרדיאנט דסנט. המאמר כולל תיאור של הארכיטקטורה והקוד של הרשת, וכן תוצאות מבחנים שמציגות את יעילות הרשת.
תקציר מקורי באנגליתarXiv:2410.11687v4 Announce Type: replace Abstract: In-context learning lets a sequence model adapt to a new task from examples in its input. A prominent line of work shows how self-attention can be constructed to implement gradient descent on a linear predictor fit to the in-context examples during the forward pass. State-space models (SSMs) and other linear recurrent networks (LRNNs) model sequences at linear time cost, but it is unclear how their recurrent update could carry out the same in-context gradient descent. We introduce Gradient-based Recurrent In-context Learner (GRIL), a diagonal LRNN that factorizes a supervised gradient step into a short-window cross-product write and a multiplicative readout of the next query. For linear regression, this construction accumulates the contex
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית