יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בלי-אינקאסט: תוכנית יוזמת של רכיבי מומחים

Incast-Free MoE Rate-Based Scheduling
אנו מציגים תוכנית יוזמת חדשה לרכיבי מומחים, שמטרתה למנוע תקלות עקב עומס חומרה. התוכנית, שנקראת MoE Rate-Based Scheduling, מסוגלת למנוע תקלות ולשמור על תפוקה גבוהה.
תקציר מקורי באנגליתarXiv:2607.26340v1 Announce Type: cross Abstract: Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. We propose an alternative proactive fair scheduling framework tailored for MoE workloads, which effectively prevents fabric oversubscription. We also outline how it can be implemented in NICs. Finally, through extensive simulations with real and synthetic workloads, we demonstrate that this framework consistently eliminates incast, maintains a near-100% link utilization, and reduces Collective Completion Time (CCT).
קרא במקור המקורי