כתבה
arXiv cs.AI ·
VACE: תצפית-מוגבלת של תהליך חלופי לשיפור דגמי סוכנים ומסגרות
VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
VACE מציג תהליך חלופי לשיפור דגמי סוכנים ומסגרות. התהליך מציע שיפור בתצורה וביצועים.
תקציר מקורי באנגליתarXiv:2609.37105v1 Announce Type: cross Abstract: Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean pa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית