כתבה
arXiv cs.AI ·
UniGuardian: הגנה מאוחדת נגד התקפות
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
UniGuardian היא שיטה לזיהוי התקפות על מודלי שפה גדולים. היא מגלה התקפות הזרקת פרומפט, התקפות backdoor והתקפות adversarial. השיטה פועלת ללא צורך באימון מחדש של המודל.
תקציר מקורי באנגליתarXiv:2502.13141v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt per
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית