כתבה
arXiv cs.AI ·
הסכנה המשוערת: כלי להגנה נגד התקפות של סוכני שפה
Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
מערכת חדשה לזיהוי התקפות של סוכני שפה שמעבר לפני שהן קורות. המערכת, הקרויה SSH, משתמשת בסימולציות של סוכני שפה קטנים כדי לבנות עץ טרסורי של תנועות עתידיות. המערכת ניתנת להתקנה כתוספת למערכות זיהוי קיימות.
תקציר מקורי באנגליתarXiv:2609.39549v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית