כתבה
arXiv cs.LG ·
הקצאת קרדיט גמישה למידה ללימודי חיזוק ל-LLM
Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning
GACA היא שיטה חדשה להקצאת קרדיט בלימודי חיזוק למודלים LLM. היא משפרת את היכולת להבין איזה חלטור השפיע על התוצאה. GACA נבדקה על ALFWorld ו-WebShop והראתה שיפור בהצלחת המשימות.
תקציר מקורי באנגליתarXiv:2609.12424v1 Announce Type: new Abstract: Reinforcement learning is now the standard way to train large language model agents on long-horizon tasks, where dozens of interdependent actions precede a single sparse reward. Critic-free, group-relative methods such as GRPO suit this regime, but they broadcast one trajectory-level scalar to every step and cannot say which decision drove the outcome. GiGPO recovers a step-level signal by grouping time steps that share an anchor state, yet it merges the step- and episode-level estimates under one fixed weight, spending the same resolution on a pivotal branching decision as on a routine, near-deterministic transition. We argue that the right resolution is state-dependent, and propose GACA, a critic-free estimator whose granularity follows an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית