כתבה
arXiv cs.AI ·
איוונט-לג'דר: תיעוד-עדות לאימות טענות
Evidence-Ledger Adjudication for Claim-Evidence Traceability
אי.אי. יכולים לכתוב טענות מהר יותר מאשר כותבים יכולים לאמת את הראיות.
תקציר מקורי באנגליתarXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית