יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

RAISED: תכנון עצמי להתנגדות להזרקה של פרומפט

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
RAISED הוא פרקטיקה של תכנון עצמי שמטרתה להגביר את ההתנגדות של סוגיות LLM להזרקה של פרומפט. הפרקטיקה נבחנה במספר סקנריות והוכיחה עצמה כיעילה.
תקציר מקורי באנגליתarXiv:2610.06401v2 Announce Type: replace-cross Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self
קרא במקור המקורי