כתבה
arXiv cs.LG ·
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
תקציר מקורי באנגליתarXiv:2609.33875v2 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית