כתבה
arXiv cs.AI ·
מסתור במבט ראשון
Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents
חוקרים פיתחו שיטה להטמעת תכנים מורעלים בסביבות LLM, על ידי הפרדה בין הטקסט המקדים לבין הפעולה עצמה. השיטה מאפשרת ליצור תכנים מורעלים שנראים לגיטימיים ומועילים, אך למעשה משרתים מטרות זדוניות. המחקר מבוסס על ניתוח של גורמי סיכון בסביבות LLM.
תקציר מקורי באנגליתarXiv:2609.39352v1 Announce Type: cross Abstract: LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית