כתבה
arXiv cs.AI ·
פער בין כוונה להתנהגות ברובוטים LLM
Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots
חוקרים גילו פער בין כוונה להתנהגות ברובוטים המבוססים על מודלי שפה גדולים. הם פיתחו כלי לבדיקת ביטחון POEF, שמטרתו לאתר פגיעויות ברובוטים אלו.
תקציר מקורי באנגליתarXiv:2412.16633v5 Announce Type: replace-cross Abstract: LLM-based robots use Large Language Models (LLMs) as planners to translate natural language instructions into policies such as grasp(), move_to(), and open_gripper(). Jailbreak attacks on these robots extend the threat from generating malicious content to executing harmful behaviors. However, we find that existing jailbreak attempts against LLM-based robots that produce malicious-looking policies (intent jailbreaks) often fail to induce harmful physical actions by robots (behavior jailbreaks), due to robot-specific constraints, such as logical errors and hallucinated control APIs. In this paper, we demystify the intent-behavior gap and investigate its root causes to inform effective defenses. Our measurement study finds that current
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית