כתבה
arXiv cs.AI ·
AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories
תקציר מקורי באנגליתarXiv:2609.06972v1 Announce Type: cross Abstract: LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read in order: a benign prefix gives way to actions that serve the attacker rather than the user. Existing benchmarks measure whether such attacks succeed against live agents, and existing guard models judge a trace as a whole; no public corpus labels, step by step, where an injection enters a trajectory and which steps it corrupts. We present AgentDrift, a benchmark of 12,536 synthetic tool-call trajectories over five agent domains in which every one of the 71,024 steps carries one of four labels: benign, inje
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית