כתבה
arXiv cs.AI ·
חוקי סקאלינג לשחרור מאסיבי של מודלי שפה גדולים: פולינומי-אקספוננטי
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
חוקי סקאלינג לשחרור מאסיבי של מודלי שפה גדולים: פולינומי-אקספוננטי. נמצא כי התקפות אדוורסריאליות יכולות להפנות מודלי שפה מאולץ להתנהגות לא בטוחה. המאמר חוקר את הסקאלינג של ההצלחה של התקפות אדוורסריאליות על מודלי שפה גדולים.
תקציר מקורי באנגליתarXiv:2603.11331v4 Announce Type: replace-cross Abstract: Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that adversarial prompt-injection attacks can amplify attack success rate from the slow polynomial growth observed without injection to exponential growth with the number of inference-time samples. We first identify a minimal statistical mechanism for these two regimes by giving a small set of assumptions on the distribution of safe generation across contexts under which both scaling laws follow. To explain this phenomenon further, we propose a theoretical generative model of proxy language in terms of a spin-glass system operating in a replica-symmetry-breaking regime, where generations are drawn from the associa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית