יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קיבוצי-אקספרטים ושיווק במודלי תרגומי חלקיקים

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers
מודלי תרגומי חלקיקים עם קיבוצי-אקספרטים משפרים את דיוק המודלים עם קיבוצי-אקספרטים. המחקר חוקר את יכולתם של מודלי תרגומי חלקיקים עם קיבוצי-אקספרטים לשפר את דיוקם. התוצאות מצביעות על כך שהמודלים עם קיבוצי-אקספרטים משפרים את דיוקם כאשר מגדילים את קיבוצי-אקספרטים. המחקר גם חוקר את השפעת קיבוצי-אקספרטים על דיוק המודלים.
תקציר מקורי באנגליתarXiv:2610.02701v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly
קרא במקור המקורי