יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

פעולות בטוחות לבדן אינן מבטיחות סוכנים בטוחים

Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
דו״ח חדש מראה כי מודלי שמירה אינם מספיקים לבטיחות סוכנים. המחקר מציג את ObligationBench, בנך„מ הראשון להערכת יכולת זיהוי חובות, ומפתח את ObligationGuard, מודל המשיג תוצאות טובות יותר.
תקציר מקורי באנגליתarXiv:2610.11773v1 Announce Type: new Abstract: Guard models are increasingly used to safeguard LLM-based agents, primarily by identifying actions that agents are forbidden to perform. However, identifying forbidden actions alone is insufficient to ensure agent safety. In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations. Our preliminary study on a popular benchmark for evaluating safety shows that 56.92% of GLM-5.3 trajectories contain unfulfilled obligations, compared with only 30.00% containing forbidden actions. This finding reveals unfulfilled obligations as a major and previously overlooked source of safety risk. However, to our knowledge, no existing benchmark evaluates whether guard models
קרא במקור המקורי