יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מה לומדים רשתות רפורמנטליות רב-מסלוליות? הטבעת קרדיט וניידות בסוכני קוד

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
רשתות רפורמנטליות רב-מסלוליות לומדות להטביע קרדיט ולהתאים לאקראיות חדשות. המחקר חוקר את תהליך הלמידה של רשתות רפורמנטליות רב-מסלוליות, ומציע דרך להטביע קרדיט ולהתאים לאקראיות חדשות. המחקר גם חוקר את יעילות הרשתות בסביבות שונות.
תקציר מקורי באנגליתarXiv:2609.04518v1 Announce Type: new Abstract: Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and comparing their rewards inside one relative-advantage group. We isolate the second choice in repository-level coding. From one Qwen3-8B supervised warm start we replay the same frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent, with the same number of updates, under two rules for group-relative policy optimization (GRPO), Within (one group per task-harness pair) and Cross (harnesses pooled within a task), and score every checkpoint with a sealed SWE-bench Verified oracle on four source harnesses and a minimal harness held out of training. The
קרא במקור המקורי