כתבה
arXiv cs.AI ·
ספסל הפיל הטרויאני: מבחן דינמי להתקפות והגנות על זיכרון קבוע באג'נטים LLM
Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents
מבחן חדש להתקפות והגנות על זיכרון קבוע באג'נטים LLM. המבחן, שנקרא ספסל הפיל הטרויאני, מדגים את יעילות התקפות והגנות על זיכרון קבוע באג'נטים LLM. המבחן כולל שני חלקים: (1) בנק אופן-אבולב, שמדגים את יעילות ההגנות והזיכרון האחורי כנגד התקפות רפויות, ו(2) בדיקת ביצועיות של תכונות הביטחון של זיכרון קבוע. המבחן נבחן על אג'נט עזר דינמי על ארבעה זיכרונות אחוריים (זיכרון כלי רפוי, זיכרון אג'נטי, RAG וחלון רצוף של זיכרון).
תקציר מקורי באנגליתarXiv:2605.01970v4 Announce Type: replace-cross Abstract: Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. The Trojan Hippo attack is a class of persistent memory attacks that operates under a more realistic threat model than prior memory poisoning work. The attacker plants a dormant payload into an agent's long-term memory via a single untrusted tool call (e.g., a crafted email), which activates only when the user later discusses sensitive topics such as finance, health, or identity, and exfiltrates high-value personal data to the attacker. While anecdotal demonstrations of such attacks have appeared against deployed systems, no prior work systematically evaluates them across heterogeneous memory arc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית