כתבה
arXiv cs.CL ·
Stepped MoE: רוטינג ברמה-פרקטיקה עם רגישות לתפוקה
Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
מאמר חדש מציג פרקטיקה של רוטינג ברמה-פרקטיקה, המאפשרת למודלים להתאים לתפוקה ולמשאבים שונים. הפרקטיקה, הנקראת Stepped MoE, משלבת רכיבי רשת רזורבייטים עם רשתות רזורבייטים רגישות לתפוקה. המאמר מציג תוצאות ניסויים המציגות את יעילות הפרקטיקה בהשוואה למודלים רגילים.
תקציר מקורי באנגליתarXiv:2610.07348v1 Announce Type: cross Abstract: Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that cond
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית