כתבה
arXiv cs.AI ·
ללמוד להחליט, לא לנמק
Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steering
חוקרים מציגים שיטה חדשה להכשרת מפעילי קבלת החלטות במודלים שפה, המאפשרת הכנסת יכולות חדשות עם פחות פרמטרים. השיטה משתמשת בהינדוס דרגה נמוכה והיא יעילה יותר משיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.06950v1 Announce Type: cross Abstract: Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce \method{}, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, wit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית