יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

סיכוני ציונים, ראיות נרדפות והצהרות מבוטלות: סקירה-מעקב של QA חזיתי

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
מודלי חזיתיים נתקלים בראיות נרדפות, ומציגים הצהרות מבוטלות עם מספרים נכונים. המאמר דן בסקירה-מעקב של QA חזיתי, ומציע שיטות לשיפור האמינות של המודלים.
תקציר מקורי באנגליתarXiv:2609.15319v1 Announce Type: cross Abstract: Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced accuracy, increased forced declarations, increased tool calls, and increased cost per correct answer. Confidence and benchmark calibration did not fully capture wrong answers; a documented production incident shows fabricated structural claims can be mixed with accurate numeric tables. Agentic evaluations need claim-level receipts (statement-level provenance, not answer-level scores), condition-aware scoring, and human-adversarial verification - an auditing discipline, not a leaderboard. The setting we measure is financial due diligence; the setting we are building toward next is defense staff w
קרא במקור המקורי