יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קיבולת מותנית וניתוב במודלים MoE

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers
חוקרים בדקו את השימוש במודלים MoE בתחום הפיזיקה. הם מצאו שניתוב מומחים משפר את הדיוק, אך לא בהכרח מוסיף עומס חישובי. המחקר פורסם ב-arXiv וכולל קוד ותצורות ניסוי.
תקציר מקורי באנגליתarXiv:2610.02701v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly asso
קרא במקור המקורי