כתבה
arXiv cs.LG ·
Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
תקציר מקורי באנגליתarXiv:2610.00400v1 Announce Type: new Abstract: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings mot
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית