כתבה
arXiv cs.AI ·
עלות הידיעה: פרוטוקול מודע למשאבים לבדיקת הזיוף מעבר לדירוג קבוע
The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
במאמר זה, פותחים פרוטוקול חדש לבדיקת הזיוף במודלי AI, המצרף את עלות החישוב. ניתן לקרוא על כך במאמר arXiv:2607.24063v2.
תקציר מקורי באנגליתarXiv:2607.24063v2 Announce Type: replace Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw factuality score (H-Score 0.9169 vs 0.9103) and would top a static leaderboard, but once cost is counted it is the worse system, losing on Q-Score (0.5169 vs 0.5217) at roughly four times the tokens and latency, under a reported cost weight whose sensitivity we sweep. So the system that tops a static leaderb
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית