יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

איך וריפייקטור יכול לחשוף את התשובה

A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging
חוקרים גילו כי וריפייקטור יכול לחשוף את התשובה בזמן ניפוי שגיאות בסביבת סימולציה. המחקר מציע שיטה חדשה לוודא שהווריפייקטור לא מחשוף את התשובה.
תקציר מקורי באנגליתarXiv:2610.00126v1 Announce Type: cross Abstract: Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the target identity, an exact optimizer may appear effective without resolving any genuine ambiguity. We report such a failure in an aggregate-trace debugger for a closed-loop decision agent. Exact minimum hitting set (MHS) and a propagation-aware greedy method returned identical supports in 12/12 development cases and the same planted-fault recovery in 9/12. A subsequent audit found that exact-anchor predicates produced the planted pair in 9/9 cases. After removing those anchors, overall planted-pair recovery was 8/9; h
קרא במקור המקורי