כתבה
arXiv cs.CL ·
מודלי שפה מורכבים של מומחים יכולים להיות חזקים וכלכליים
Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers
מודלי שפה מורכבים של מומחים מבצעים השוואה טובה יותר למודלי רוכבים צפופים עם כמות פרמטרים פעילה דומה.
תקציר מקורי באנגליתarXiv:2609.13486v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong first-stage retrievers, with effectiveness improving as backbones grow in size. However, every query and document must pass through the full model, so encoding cost increases with model size. Mixture-of-Experts (MoE) LLMs activate only a subset of parameters per token and are widely used to scale generative models, yet remain underexplored as retrievers. We systematically study MoE backbones for retrieval by training MoE and dense LLMs from several families using the same procedure, evaluating them across diverse datasets, and measuring query encoding time under the same serving configuration. We show that MoE retrievers outperform d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית