כתבה
arXiv cs.AI ·
ToolFence: תיקון זכויות דק-משקל לאג'נטים LLM עם כלי
ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
ToolFence מציע תיקון זכויות דק-משקל לאג'נטים LLM עם כלי, כדי למנוע פגיעה עקיפה. התיקון נבחן על AgentDojo עם Qwen3-max והציג תוצאות מוצלחות.
תקציר מקורי באנגליתarXiv:2609.37196v1 Announce Type: cross Abstract: Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and whe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית