יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ESTS ב-WMT26: גזירת מומחים מודרכת נתיבים לדחיסת מודל

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression
ESTS הגישה שישה מודלים ל-WMT26, כולל GPT-OSS-20B. המודלים עברו גזירת מומחים ודחיסה. התוצאות הראו שיפור בביצועים.
תקציר מקורי באנגליתarXiv:2609.12310v1 Announce Type: new Abstract: We describe six submissions under the team name ESTS to the unconstrained WMT26 Model Compression Shared Task for English--Simplified Chinese and English--Egyptian Arabic. We submit three compression operating points per translation direction, all derived from GPT-OSS-20B. We use task-specific routing mass to rank experts and cross-lingual routing divergence to allocate retained capacity across layers, then physically remove low-importance experts. The resulting specialists are recovery-tuned on GPT-5.1-generated synthetic translation data and further compressed by applying MXFP4 quantization to the retained expert projection weights. We additionally implement a robust inference system for the instruction-conditioned WMT26 setting, including
קרא במקור המקורי