כתבה
arXiv cs.CL ·
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
תקציר מקורי באנגליתarXiv:2607.24062v1 Announce Type: cross Abstract: Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separation between training and inference engines, and the use of low-precision quantization in inference versus higher-precision computation in training. To address training instability issues caused by high training-inference discrepancy, we present the principles and methods for its adaptive control. We propose Adaptive Control Reinforcement Learning (ACRL), which adaptively maintains the training-inference discrepancy within a reasonable range to ensure stable RL training. Beyond stabilization, ACRL
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית