כתבה
arXiv cs.AI ·
מצא את המוצא הנכון: התנהגות של דגימות-מודל והתאמה למשימות של סוכנים
Finding the Right Fit: Model-Harness Interactions across Agent Tasks
התנהגות של דגימות-מודל והתאמה למשימות של סוכנים משפיעה על יכולתם לבצע משימות. נמצא כי דגימות-מודל והתאמה למשימות של סוכנים יכולים להיות חשובים להצלחתם של סוכנים. המחקר חשף כי דגימות-מודל והתאמה למשימות של סוכנים יכולים להיות חשובים להצלחתם של סוכנים.
תקציר מקורי באנגליתarXiv:2610.00917v1 Announce Type: new Abstract: Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.6
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית