יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אבחון Training-Inference Mismatch ב-LLM RL

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference
חוקרים פיתחו שיטה לאבחון Training-Inference Mismatch ב-LLM RL. המחקר מראה כי הפרשים קטנים ברמת הטוקנים יכולים לגרום לקריסת אימון. התוצאות מצביעות על כך ש-Training-Inference Mismatch אינו רעש נומרי תמים, אלא הפרעה ברמת המערכת.
תקציר מקורי באנגליתarXiv:2605.14220v2 Announce Type: replace-cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this work, we isolate TIM in a zero-mismatch diagnostic setting (VeXact), and show that small token-level numerical disagreements can independently cause training collapse. We further show that TIM changes the effective optimization problem, and identify a set of remedies that could mitigate TIM. Our
קרא במקור המקורי