כתבה
arXiv cs.LG ·
הפקת רגישות לאימון אפיזודי לאג'נט LLM על-זמן-ארוך
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
אימון אפיזודי לאג'נט LLM על-זמן-ארוך: הפקת רגישות והפחתת רמות גאוס
תקציר מקורי באנגליתarXiv:2511.20718v3 Announce Type: replace Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית