יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

תרגום יעיל של טרנספורמרים ל-Mamba

Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
אנו מציגים פרקטיקה של תרגום יעיל של טרנספורמרים למודלי מדף-מצב (Mamba) דרך גשר של תשומת לב.
תקציר מקורי באנגליתarXiv:2510.19266v3 Announce Type: replace Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scratch remains computationally intensive, and the ecosystem around them is far less mature than that of Transformers. Moreover, the architectural differences between SSMs and Transformers make it challenging to efficiently transfer knowledge from pretrained Transformers. In this work, we propose Cross-architecture distillation via Attention Bridge (CAB), a distillation framework that transfers attention-related representations from Transformer teachers to state-space student models. Unlike conventional knowledge distillation that supervises only final predictions, CAB enables token-level intermed
קרא במקור המקורי