כתבה
arXiv cs.AI ·
The Surface You Test Is Not the Surface That Breaks
תקציר מקורי באנגליתarXiv:2605.30454v2 Announce Type: replace-cross Abstract: Prompt-injection benchmarks for LLM agents typically test attacks through a single injection surface and report the resulting attack success rate as a property of the model. We ask whether those robustness conclusions remain stable when the same adversarial content enters through a different part of the agent interface. Using AgentDojo, we evaluate 13 LLMs across four task suites and place a byte-identical payload either in a tool output or in the tool description. This small change produces large differences in comparative robustness: 44.9% of all model pairs change their relative ordering across the two surfaces, with substantial ranking instability in every suite. The effect is especially pronounced for a small number of models,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית