כתבה
arXiv cs.AI ·
שיפור הערכת התפתחות Harness לסוכנים
Rethinking the Evaluation of Harness Evolution for Agents
חוקרים בדקו את היעילות של התפתחות Harness אוטומטית עבור סוכנים LLM. הם מצאו שהשיטה אינה עולה על שיטות פשוטות יותר ברוב המקרים, אך מראה הבטחה במשחקים ארוכי טווח.
תקציר מקורי באנגליתarXiv:2607.12227v3 Announce Type: replace Abstract: Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues in the protocol. First, prior work does not compare these approaches with simple task-level search baselines under matched feedback and inference budgets. Second, prior work searches for harness configurations using verification signals (e.g., unit test cases) drawn from the same benchmarks on which it reports the final performance of the evolved harnesses, violating the standard separation between training and test data. To address this, we compare automatic harness evolution with si
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית