יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מה החלון אינו כולל: תקיפה של ראשוניות בבסיס הביקורת של דוקומנט-גראונד

What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark
בדיקת ראשוניות בבסיס הביקורת של דוקומנט-גראונד של חוסר יציבות
תקציר מקורי באנגליתarXiv:2609.06147v1 Announce Type: cross Abstract: Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built benchmark can carry: items whose evidence is missing from the window of text the model is shown. The audit flags 36 items and separates two failures a single flag would conflate: evidence genuinely absent from the window and answers that must be computed from numbers the window does supply. Flagged items change their answers far more often, wobbling at 0.255 against 0.087 on the 427 clean items, and excluding them cuts appare
קרא במקור המקורי