כתבה
arXiv cs.LG ·
תקיפות עם פרומפטים נייטרליים
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
חוקרים גילו תקיפת פרומפטים נייטרליים שיכולה ליצור סיכונים בשרשרת האספקה של תוכנה. התקיפה משתמשת בפרומפטים תמימים כדי לגרום למודלים לייצר שמות חבילות שאינם קיימים. המחקר בדק מודלים שונים, כולל LLMs, ומצא שהתקיפה יכולה לעקוף הגנות קיימות.
תקציר מקורי באנגליתarXiv:2605.29354v2 Announce Type: replace-cross Abstract: LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package installation commands. This creates a new software supply chain risk: when an agent hallucinates a non-existent package, an attacker may register the hallucinated name and later compromise users who install it. Existing package hallucination attacks and defenses primarily focus on naturally occurring hallucinations, targeted dependency steering, or post-hoc package validation. In this paper, we introduce \emph{Neutral Prompting Attack} (NPA), a highly stealthy attack paradigm in which semantically benign instructions, such as encouraging imagination and exhaustiveness, increase packag
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית