יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הפקת רגישות לאימון רפלקסיבי סטבילי ל-LLM Agent

Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
אימון רפלקסיבי סטבילי ל-Agent LLM באורך זמן ארוך. המחברים הציגו פרקטיקה חדשה לאימון Agent LLM באורך זמן ארוך, המבוססת על טכניקות של Turn-Level Importance Sampling ו-Clipping-Triggered Normalization. הפרקטיקה, הנקראת SORL, נועדה לסטבל ולהפקיד רגישות לאימון רפלקסיבי באורך זמן ארוך. המחברים הציגו תוצאות של ניסויים שהראו כי SORL מצליח לסטבל את האימון ולהפקיד רגישות, וכי הפרקטיקה יכולה להיות יעילה במגוון של תחומים.
תקציר מקורי באנגליתarXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing
קרא במקור המקורי