כתבה
arXiv cs.LG ·
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
תקציר מקורי באנגליתarXiv:2609.35932v1 Announce Type: cross Abstract: Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית