כתבה
arXiv cs.LG ·
Forensic Trajectory Signatures for Agent Memory Poisoning Detection
תקציר מקורי באנגליתarXiv:2606.30566v2 Announce Type: replace-cross Abstract: We discover a behavioral invariant in LLM agents under persistent memory poisoning and characterize its deployment boundary. In architectures where retrieval is routed through observable memory-tool invocations, successful attacks require calling memory_recall_fact before email_send_email, a transition mechanistically forced by the attack's information-retrieval dependency. A simple rule exploiting this invariant achieves AUC = 0.9563; a Random Forest over 19 trajectory features refines it to AUC = 0.9904 (BCa 95% CI [0.987, 0.993]). The signature is overdetermined within the poisoned-but-defended evaluation set: removing all recall-related features leaves AUC unchanged. Cross-model hold-out on 9 models (7B-120B) confirms AUC = 1.00
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית