יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

SCOUT: סינרגיזציה של תפיסה ושימוש בכלי לבטיחות חישובי

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
SCOUT היא פלטפורמה של שני שלבים לבטיחות חישובי, המשלבת תפיסה רפרנטיבית עם חיפוש כלי. היא נועדה לזהות סכנות חישוביות ולצמצם את הסיכון של חישובי חסרי תקינה. SCOUT כבר הוכיחה את עצמה במבחנים שונים, כולל AutoElicit-Bench ו-OS-Blind.
תקציר מקורי באנגליתarXiv:2609.36201v1 Announce Type: new Abstract: Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verif
קרא במקור המקורי