כתבה
arXiv cs.AI ·
שיפור תגמולים תלויי-תלות ללמידה רפורמינטלית
Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
DARS (Dependency-Aware Reward Shaping) - שיטה לשיפור תגמולים בלמידה רפורמינטלית. השיטה מייצגת קדימות במשימה כפרדיקטים שקשורים על ידי יחסי תלות. DARS משתמשת בגרף של תלות כדי ליישם תגמולים על-פי-צעד. השיטה נבחנה ב-5 משפחות של משימות ובמודלים של 1.5 מיליארד עד 8 מיליארד תווים.
תקציר מקורי באנגליתarXiv:2610.01207v1 Announce Type: new Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest brok
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית