יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מרכזיות להפרדה וחזרה: רוטינג דרגה יעילה בקבוצות סיבוכיות של MoE

From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
קבוצות סיבוכיות של MoE מציגות דרגה יעילה של רוטינג המתאפיינת במסלול נמוך-גבוה-נמוך רפלקסיבי.
תקציר מקורי באנגליתarXiv:2609.06403v1 Announce Type: new Abstract: Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout graph built from MoE expert-routing similarity. Across ten MoE configurations and five math/science benchmarks, deff exhibits a reproducible low-high-low trajectory, with a prominent interior maximum in 98.5% of 3,105 model-question cohorts: routing similarity is concentrated early, maximally differentiated at intermediate budgets, and reconcentrated later, and the timing of this maximum varies systematically with architecture and reasoning effort. An exact decomposition separate
קרא במקור המקורי