כתבה
arXiv cs.LG ·
FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
תקציר מקורי באנגליתarXiv:2610.07898v1 Announce Type: new Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next traj
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית