כתבה
arXiv cs.AI ·
SHARPO: שיטה חדשה להקצאת אשמה בלמידה חיזוקית
SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
SHARPO היא שיטה חדשה להקצאת אשמה בלמידה חיזוקית, המשפרת את יכולת הלמידה של מודלים גדולים. היא משתמשת ב-Qwen2.5-7B-Instruct ומציגה תוצאות טובות יותר מאשר שיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.00838v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית