יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הסטרוקטורה של בחירת המומחים של MoE ללמידת רפלקסיה

Structuring MoE Expert Selection for Agentic Reinforcement Learning
במאמר זה, נחקרה הקשר בין התנהגות של סוכנים ובחירת מומחים של MoE. נציגה פרקטית של רשתות MoE, נצפה כי בחירת המומחים נוטה להתאים למסלולי הסוכן. נציגה פרקטית של רשתות MoE, נצפה כי בחירת המומחים נוטה להתאים למסלולי הסוכן. נציגה פרקטית של רשתות MoE, נצפה כי בחירת המומחים נוטה להתאים למסלולי הסוכן.
תקציר מקורי באנגליתarXiv:2610.07332v1 Announce Type: cross Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically lim
קרא במקור המקורי