יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בחינת סוכנים במסגרת חוזים רציפים: כאשר תחבולה עלולה לפגוע בעקביות או באיכות

Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality
במאמר זה, נחקרה השפעת חוזים רציפים על יכולתם של סוכנים לפעול. נמצא כי תחבולה בחוזים עלולה לפגוע בעקביות או באיכות. המחקר נערך על ידי צוות של CodeAct, והתפרסם בarXiv.
תקציר מקורי באנגליתarXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task progress. Training traces demonstrate task-solving strategies and runtime-specific ways to store and recover intermediate results. We study runtime transfer: whether agents trained under one contract remain effective under the other, a dependence that fixed-runtime evaluations can conceal. Across three tasks requiring a working record built from tool feedback, we fine-tune separate Qwen3-8B agents per task and runtime on instance-paired persistent and stateless traces and evaluate all four training-deployment combina
קרא במקור המקורי