כתבה
arXiv cs.LG ·
למידת חיזוק עם משימות מפורקות
Reinforcement Learning with Decomposed Subtasks
חוקרים מציגים שיטה חדשה ללמידת חיזוק, המפרקת משימות לתת-משימות. השיטה, הנקראת RLDS, מאפשרת למודלים ללמוד מהתנסות בצורה יותר יעילה. החוקרים בדקו את השיטה בארבעה מבחנים שונים וגילו שהיא משפרת את הביצועים, במיוחד במשימות עם הטרוגניות גבוהה.
תקציר מקורי באנגליתarXiv:2609.27035v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית