יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

FinalityBench: בנקאי רשמי לבדיקת החלטות סוכנים תחת דחיות וסתירות כספיות

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
בנקאי רשמי לבדיקת החלטות סוכנים תחת דחיות וסתירות כספיות. הבנקאי נועד לבדוק את החלטות סוכנים במצבים שבהם יש דחיות וסתירות בין המערכות השונות. הבנקאי כולל 321 משימות, כולל 45 זוגות תאומים (90 משימות) שבהם המערכות השונות חולקות תצפיות זהות ברגע ההחלטה, ובסופו של דבר יש פעולות נכונות שונות. הבנקאי נועד לבדוק את היכולת של סוכנים להחליט במצבים שבהם יש דחיות וסתירות.
תקציר מקורי באנגליתarXiv:2609.04706v1 Announce Type: new Abstract: A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated, dropped and reordered, so for minutes at a time the four hold contradictory beliefs about the same order. An agent resolving the exception must decide whether to ship goods, re-submit a capture, refund or wait, knowing some of those cannot be undone. We present FinalityBench, an executable benchmark for that decision. It keeps a hidden canonical event log and derives each system's view from a separately faulted delivery stream, so disagreement follows from specified fault semantics rather than being authored. Grading is on executed monetary effects: an episode is scored by the merchant's terminal economic position, relative to a pri
קרא במקור המקורי