כתבה
arXiv cs.AI ·
הפחתת סבילות בהכשרת LLM לטורים ארוכים
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
במאמר זה, צוות מדעניות ומדעני חומרה מציגים פתרון לבעיית הסבילות בהכשרת LLM לטורים ארוכים. הם מציגים את SORL, פרקטיקה של הכשרה של LLM לטורים ארוכים, שמטרתה להפחית את הסבילות בהכשרה. הם מציגים גם שני אלגוריתמים, SO-PPO ו-SO-GRPO, שמבוססים על SORL. האלגוריתמים הללו נבדקו על ידי הצוות במספר תחומים, כולל QA, QA רב-קפיצה, ו- QA רפואית. התוצאות הראו כי SORL מספק פרקטיקה יעילה וקצרת-זמן להפחתת סבילות בהכשרת LLM לטורים ארוכים.
תקציר מקורי באנגליתarXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית