יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

TRACE: אימון רב-דיוק למודלי שפה

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
TRACE הוא כלי לאימון יעיל של מודלי שפה מסוג Mixture-of-Experts. הוא משפר את דיוק האימון ומקטין את זמן האימון. TRACE משתמש בשיטות קוונטיזציה מתקדמות כדי לשפר את ביצועי המודל.
תקציר מקורי באנגליתarXiv:2610.07767v1 Announce Type: cross Abstract: Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization
קרא במקור המקורי