יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ניווט מותאם לירידה: רווחה למומחים LoRA משולבים

Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts
חוקרים פיתחו שיטה חדשה לניווט מומחים LoRA, המבוססת על רווחה לירידה. השיטה, הנקראת VANE, משפרת את הביצועים של מודלים כגון Llama-3.2-3B ו-Llama-3.1-8B, עם פחות פרמטרים אימונים. VANE מנצלת את העובדה שהרווחה של כל זוג מומחים תלויה במכפלה הפנימית בין העדכון לבין הגרדיאנט השלילי של ההפסד.
תקציר מקורי באנגליתarXiv:2610.00493v1 Announce Type: new Abstract: Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations nev
קרא במקור המקורי