כתבה
arXiv cs.LG ·
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
תקציר מקורי באנגליתarXiv:2610.11775v1 Announce Type: cross Abstract: Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most pre
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית