יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

TRACE: גידור-מנהיגי קוונטיזציה-מודע לאימון לטיפול ב-FP4 של מודלי LLM של MoE

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
TRACE היא תשתית קוונטיזציה-מודעה ל-FP4 לאימון RL של מודלי LLM של MoE. התשתית משתמשת בגידור-מנהיגי קוונטיזציה-מודע לאימון RL של MoE. TRACE מאפשרת אימון RL של MoE בעל FP4 עם קצב רולאוט של 5.4x.
תקציר מקורי באנגליתarXiv:2610.07767v1 Announce Type: new Abstract: Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-a
קרא במקור המקורי