יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ציוני נקיים, ראיות דולקות והכרה חד-צדדית: תיקון-קבלה של חיפוש נושאי QA

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
מודלי גבול נכשלים בעתירות, וזה גורם להכרה חד-צדדית בתשובות שגויות. נדרשים תיקון-קבלה ובדיקה נוקבת.
תקציר מקורי באנגליתarXiv:2609.15319v1 Announce Type: cross Abstract: Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced accuracy, increased forced declarations, increased tool calls, and increased cost per correct answer. Confidence and benchmark calibration did not fully capture wrong answers; a documented production incident shows fabricated structural claims can be mixed with accurate numeric tables. Agentic evaluations need claim-level receipts (statement-level provenance, not answer-level scores), condition-aware scoring, and human-adversarial verification - an auditing discipline, not a leaderboard. The setting we measure is financial due diligence; the setting we are building toward next is defense staff w
קרא במקור המקורי