יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TwinCheck: וריפיקציה של תאומים שליליים לאימות של כלי כלי

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents
TwinCheck היא פוליסה לאימות בזמן הריצה של כלי כלי, המבצעת וריפיקציה של תאומים שליליים. היא מבנה תאומים שליליים, ומחליפה את ההצעה של הכלי רק אם התאום מצליח בבדיקות סטרוקטורליות והמערכת מעדיפה אותו.
תקציר מקורי באנגליתarXiv:2609.26911v2 Announce Type: replace Abstract: A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted rep
קרא במקור המקורי