כתבה
arXiv cs.AI ·
הפניה נגד-טבעית: ריפליי עם סביבות ניתנות להפרדה כפרסומים חינם לאגנטים של תכנות
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
אגנטים לתכנות: פרוצדורה חדשה של ריפליי עם סביבות ניתנות להפרדה, שמשפרת את הביצועים של אגנטים לתכנות. הפרוצדורה, שנקראת Counterfactual Rollout Replay (CRR), משתמשת בסביבות ניתנות להפרדה כדי לקבל חזרה על צעדי האגנט. CRR משפרת את הביצועים של האגנטים בכ-5% בהשוואה לאגנטים שלא השתמשו ב-CRR. הפרוצדורה נחשבת לחדשה ומשמעותית, ומציעה דרך חדשה לאימון אגנטים לתכנות.
תקציר מקורי באנגליתarXiv:2609.33875v2 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית