כתבה
arXiv cs.LG ·
אמינות הערכת סוכנים: משימות נוספות לא תמיד יכולות לתקן
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
חוקרים בדקו את אמינות הערכת סוכנים ומצאו כי הוספת משימות נוספות לא תמיד משפרת את הדיוק. הם פיתחו מסגרת בייסיאנית לניתוח האמינות ויישמו אותה על 22 בנצ'מרקים. התוצאות מראות כי אמינות הערכה תלויה במטרת המדידה וכי בחירת הסקאפולד יכולה לשנות את המסקנות.
תקציר מקורי באנגליתarXiv:2610.00651v1 Announce Type: cross Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would improve them. We develop a Bayesian variance-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index. The framework separates signal, performance differences relevant to the intended claim, from noise, irrele
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית