כתבה
arXiv cs.CL ·
$S^3$: שיפור יעילות מודלי תיקוף
$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
חוקרים הצליחו לשפר יעילות מודלי תיקוף באמצעות שיטה חדשה הנקראת $S^3$. השיטה מאפשרת להפחית את כמות הטוקנים הנדרשים לתיקוף, תוך שמירה על דיוק. השיטה נבדקה על מודלים שונים, כולל LLM, והראתה שיפורים משמעותיים.
תקציר מקורי באנגליתarXiv:2609.37976v1 Announce Type: cross Abstract: LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית