כתבה
arXiv cs.LG ·
Position: Evaluation Scores Are Perishable Knowledge Claims
תקציר מקורי באנגליתarXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית