יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ParanoiaEval: בדיקת עבודה מגננת מיותרת בקידוד אגנטי

ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
ParanoiaEval הוא בנץ'מארק ראשון מסוגו להערכת יכולות טיפול בסיכונים בסוכנים מקודדים. הוא מאפשר לבדוק האם סוכנים מקודדים עוסקים בעבודה מגננת מיותרת. המחקר מראה כי 11.2%-58.7% מהריצות כוללות עבודה מגננת מיותרת, למרות ראיות מפורשות.
תקציר מקורי באנגליתarXiv:2610.08662v1 Announce Type: new Abstract: As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We
קרא במקור המקורי