כתבה
arXiv cs.LG ·
פציעה על פי תווית: כללי זיהוי פגמים-מלאכותיים בשטח
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
במאמר זה, המחברים חוקרים את יעילות זיהוי פגמים-מלאכותיים ב-LLMs, ומציעים פתרון זול לשיפור זיהוי פגמים-מלאכותיים. הם מציעים להוסיף תווית סיווג לאחר הפנייה של המשתמש, כדי לשפר את זיהוי הפגמים-מלאכותיים. המחברים חוקרים את יעילות פתרון זה במספר תרחישים, ומציעים פתרון זול לשיפור זיהוי פגמים-מלאכותיים.
תקציר מקורי באנגליתarXiv:2610.02413v1 Announce Type: new Abstract: LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classification instruction after the user's turn and read the probe at that point, to sharpen it: the instruction asks the model to represent the incoming request as a class, concentrating the signal the probe must separate, at negligible serving cost. But does the wording of that suffix matter, and does its benefit hold in the wild, on attack types the probe never saw in training, the regime a deployed monitor faces? We test this with a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית