יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

T-Router: חידוש בתפנות תלמית לשכנוע עם רכיבת רפורמנט עמידה-פרמטרים

T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning
מאמר חדש עוסק בשכנוע עם רכיבת רפורמנט עמידה-פרמטרים. ה-T-Router, חידוש זה, מציע תפנות תלמית שמתמקדת בשימוש מחדש של חישובים כבר נעשו. זה יכול להפחית את כמות הפרמטרים הדרושים ולשפר את יעילות הלמידה.
תקציר מקורי באנגליתarXiv:2609.39109v1 Announce Type: cross Abstract: Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achie
קרא במקור המקורי