יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הפרגיליות של מנגנוני תגים-מסווגים לזיהוי שימוש לא רצוי ב-LLMs פתוח-משקל

The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
מנגנוני תגים-מסווגים לזיהוי שימוש לא רצוי ב-LLMs פתוח-משקל נמצאו לא יציבים. חוקרים גילו שאפשר להפיל אותם בקלות. זאת על פי מאמר שפורסם ב-arXiv.
תקציר מקורי באנגליתarXiv:2610.03124v1 Announce Type: cross Abstract: Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emph{token-level trigger-tags}, which introduce watermark
קרא במקור המקורי