כתבה
arXiv cs.AI ·
מה החלון אינו כולל: אודיט של נכסים בבסיס תעודה של פגיעה
What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark
אודיט של נכסים בבסיס תעודה של פגיעה, שבו נבדקה יציבות של דגימות
תקציר מקורי באנגליתarXiv:2609.06147v1 Announce Type: cross Abstract: Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built benchmark can carry: items whose evidence is missing from the window of text the model is shown. The audit flags 36 items and separates two failures a single flag would conflate: evidence genuinely absent from the window and answers that must be computed from numbers the window does supply. Flagged items change their answers far more often, wobbling at 0.255 against 0.087 on the 427 clean items, and excluding them cuts appare
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית