יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מה עובדה כשמדדים רשימת ראשי רשימות LLM?

What Does an LLM-Agent Leaderboard Rank Actually Compare?
מדדי רשימת ראשי רשימות LLM לא תמיד תומכו בטענות של עליונות. ניתוח חדש מצביע על חשיבות ניתוח האסטימנד והביטחון.
תקציר מקורי באנגליתarXiv:2609.07785v1 Announce Type: new Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBe
קרא במקור המקורי