כתבה
arXiv cs.LG ·
TRIAGE: סטביליזציה של מניעת תיאום ב-RF4
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
TRIAGE - שיטה לסטביליזציה של מניעת תיאום ב-RF4, שמשפרת את אופטימיזציה של רכיבי RL בעזרת ניתוח ותיקון של תיאום.
תקציר מקורי באנגליתarXiv:2610.07043v1 Announce Type: new Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. Their tail tokens become concentrated in a small fraction of response segments before mismatch spreads globally. Motivated by these findings, we introduce TRIAGE, a direction-aware stabilization
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית