כתבה
arXiv cs.LG ·
העברה ב间-בנצ'מרקים מ-RL על משימות קוד אגנטי
Cross-Benchmark Transfer from RL on Agentic Coding Tasks
חוקרים בדקו האם למידת חיזוק (RL) על משימות קוד אגנטי יכולה לשפר את היכולת של מודלים לבנות תכונות. הם השתמשו במודל Kimi K2.7 Code וראו שיפור משמעותי בביצועים.
תקציר מקורי באנגליתarXiv:2610.00890v1 Announce Type: new Abstract: Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית