יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Stepped MoE: רוטינג פחות-סגור עם רמת סבילות ניתנת להגדרה

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
מאמר חדש: Stepped MoE - פלטפורמה מאוחדת למודלי שפה ניתנים להתאמה. המאמר עוסק בפיתוח מודל שמתאים למגוון סצנות הפריצה, כולל גם פריצה לגביים. המודל משתמש בארכיטקטורה של MoE (Multi-Headed Expert) ובכך הוא יכול להתאים למגוון רמות סבילות. המאמר כולל ניסויים שמדגימים את יעילות המודל.
תקציר מקורי באנגליתarXiv:2610.07348v1 Announce Type: new Abstract: Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that condit
קרא במקור המקורי