כתבה
arXiv cs.CL ·
פרצות מדיניות בהערכת סוכנים
Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error
חוקרים גילו פרצות מדיניות בהערכת סוכנים. מדיניות לא ברורה יכולה לגרום לתוצאות לא אמינות. יש צורך לבדוק מדיניות לפני איסוף נתונים.
תקציר מקורי באנגליתarXiv:2609.14400v1 Announce Type: new Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $\tau^2$-bench domains, we develop a taxonomy of such policy loopholes and show that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials. A cross-domain comparison reveals that exploitability requires both policy ambiguity and tool permissiveness: when policy complexity exceeds what tools can enforce, agents resolve gaps inconsistently and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית