יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

RAISED: תרגום-עצמי להתמודדות עם פצצות-פרומפט

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
RAISED הוא תרגום-עצמי שמטפל בפצצות-פרומפט ב-LLM. הוא משתמש בעצמו-התאמה ועצמ-התפשטות להגן על עצמו. זה עוזר ל-LLM להיות יותר חזק ולא להיות פגיע לפצצות-פרומפט.
תקציר מקורי באנגליתarXiv:2610.06401v2 Announce Type: replace-cross Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self
קרא במקור המקורי