כתבה
arXiv cs.LG ·
פריצת כלא: פרקטיקה מידע-תאורטית להתקפות חידושי-קומפוזיציוניות
Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks
במאמר זה, המחברים פיתחו פרקטיקה חדשה להתקפות חידושי-קומפוזיציוניות ב-LLMs. הם חקרו את האפשרות להסתיר את המטרה הרעה באמצעות חיבור של תפקידים חיוביים. הם פיתחו פרקטיקה מידע-תאורטית להתקפות חידושי-קומפוזיציוניות, שמאפשרת לבחור תפקידים חיוביים שיסתירו את המטרה הרעה.
תקציר מקורי באנגליתarXiv:2610.02302v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositional intent-hiding jailbreaks from an information-theoretic perspective. Our formulation associates each task with an estimated probability of being judged harmful: the average over the full task collection defines the prior probability of harmful intent, while the average over a selected bundle containing the target defines the posterior. Selecting auxiliary tasks so that these averages agree, which we call prior-posterior ma
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית