כתבה
arXiv cs.AI ·
ביקורת על יציבות השקלול של סוכני AI: יותר משימות לא תפתור (בכל פנים) את הבעיה
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
במאמר זה נבחן את יציבות השקלול של סוכני AI. נמצא כי יותר משימות לא תפתור את הבעיה של יציבות השקלול. המחברים מציעים פתרון לבעיה זו.
תקציר מקורי באנגליתarXiv:2610.00651v1 Announce Type: new Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would improve them. We develop a Bayesian variance-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index. The framework separates signal, performance differences relevant to the intended claim, from noise, irreleva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית