יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת חיזוק עם משימות מפורקות

Reinforcement Learning with Decomposed Subtasks
חוקרים מציגים שיטה חדשה ללמידת חיזוק, המפרקת משימות לתת-משימות. השיטה, הנקראת RLDS, מאפשרת למודלים ללמוד מהתנסות בצורה יותר יעילה. החוקרים בדקו את השיטה בארבעה מבחנים שונים וגילו שהיא משפרת את הביצועים, במיוחד במשימות עם הטרוגניות גבוהה.
תקציר מקורי באנגליתarXiv:2609.27035v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (
קרא במקור המקורי