יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בקרת סטריאוטיפים מוגבלת במודלים של MoE

Limited Stereotype Control Through Routing Reweighting in MoE Language Models
חוקרים בדקו את היכולת לשלוט בסטריאוטיפים במודלים של MoE. הם הציגו את FARE, כלי אבחון שמשלב פרופילים דמוגרפיים ובחירת שכבות. הניסויים הראו שינויים קטנים בהעדפות ובמדדים אחרים.
תקציר מקורי באנגליתarXiv:2603.27141v2 Announce Type: replace Abstract: Demographic prompts are routed differently from neutral prompts in Mixture-of-Experts (MoE) language models, motivating tests of routing-level stereotype control. We introduce Fairness-Aware Routing Equilibrium (FARE), a diagnostic framework combining demographic routing profiles, empirical layer selection, and fixed inference-time reweighting, and evaluate five MoE architectures in English. At the selected operating points, CrowS-Pairs preference changes by at most 1.3 percentage points; DeepSeekMoE selects no intervention. Paired 95% confidence intervals exclude decreases larger than 2.2 points on each intervened model, and the only nominally significant change (Qwen1.5, p = 0.015) does not survive multiple-comparison correction. OLMoE
קרא במקור המקורי