יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בדיקת רשת, ולא שכבה: האם דלתות סטטיסטיות ו-LLM לפעולות גורמי פעולה נכשלים בעצמן?

Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?
במחקר זה, נבדקה האם דלתות סטטיסטיות ו-LLM לפעולות גורמי פעולה נכשלים בעצמן. התוצאות הראו שהדלתות נכשלו בעצמן, ושהקשר ביניהן היה חשוב.
תקציר מקורי באנגליתarXiv:2610.07359v1 Announce Type: new Abstract: Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers ({\phi} median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers ({\phi} median +0.014, 0 of 4 significant, floors
קרא במקור המקורי