כתבה
arXiv cs.LG ·
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
תקציר מקורי באנגליתarXiv:2608.23873v3 Announce Type: replace-cross Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and dangerous actions. Adding a non-textual channel to the model's input -- a way to communicate span identity beyond text -- mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an ove
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית