יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

העברת קצב הלמידה לארכיטקטורות תרמוסטורמי-SSM

Learning Rate Transfer for Hybrid Transformer-SSM Architectures
במאמר זה, המחברים חוקרים את העברת קצב הלמידה לארכיטקטורות תרמוסטורמי-SSM. הם חוקרים את הפער בין החוקים התאורטיים לבין הimplementations הפרקטיות של hybrid architectures.
תקציר מקורי באנגליתarXiv:2610.01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original $\mu$P prescription, even though SSM operations fall outside its Tensor Programs representability conditions
קרא במקור המקורי