כתבה
arXiv cs.AI ·
SAPO: אופטימיזציה אוטומטית של פוליציה יחידה ללמידת מודלים גדולים
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
SAPO היא פרקטיקה חדשה ללמידת מודלים גדולים שמשלבת אופטימיזציה של פוליציה יחידה ולמידת ערך. היא משמשת לשפים גדולות כמו Qwen2.5-1.5B/7B ו-Qwen3-14B.
תקציר מקורי באנגליתarXiv:2608.19842v3 Announce Type: replace Abstract: Agentic reinforcement learning (RL) has emerged as an important post-training approach for enhancing the capabilities of Large Language Models (LLMs). However, existing methods face a trade-off between policy performance and resource efficiency. Conventional Proximal Policy Optimization (PPO) implementations incur substantial memory overhead from a separate critic, whereas critic-free group-relative methods require multiple rollouts and face potential learning bottlenecks on long-horizon tasks. In this work, we propose Single-rollout Autoregressive Policy Optimization (SAPO), an efficient PPO-style framework that unifies policy optimization and value learning within a single causal language model. SAPO exploits the autoregressive structur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית