יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CounterSteer: דיכוי הזרקת פקודות עקיפה

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
CounterSteer היא הגנה בזמן הרשאה שמדכאת התנהגות זרקת פקודות עקיפה. היא פועלת על ידי יצירת זרימה נגדית שמותקן בזמן הרשאה, ומורידה מכל תוצרי כלי. ההגנה זקוקה לאף תפעול נוסף, ואינה זקוקה לאימון מחדש. CounterSteer נבחנה על 5 דגמי מודלים (8B-106B) והצליחה לצמצם את קצב ההתקפה העקיפה מ-0.21-1.00 ל-0.00-0.17.
תקציר מקורי באנגליתarXiv:2609.36570v1 Announce Type: cross Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages),
קרא במקור המקורי