כתבה
arXiv cs.CL ·
RouterInterp: הבנת התמחות המשוקללת בניתוב Mixture of Experts
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
RouterInterp הוא שיטה לפירוש ניתוב מומחים במודלים Mixture of Experts. השיטה מזהה מאפיינים הקשורים לקבלת החלטות ניתוב ומייצר הסברים בשפה טבעית. היא משפרת את דיוק הזיהוי ב-65% לעומת שיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.11775v1 Announce Type: cross Abstract: Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most pre
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית