כתבה
arXiv cs.AI ·
ParanoiaEval: בדיקת פאניקה: שיפור בהערכת עבודה תקיפה בקוד
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
ParanoiaEval היא הבדיקה הראשונה של עבודה תקיפה לא נדרשת בקוד. המחקר חוקר את הקפאות סיכונים בקוד ומציע פתרונות.
תקציר מקורי באנגליתarXiv:2610.08662v2 Announce Type: replace Abstract: As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית