כתבה
arXiv cs.LG ·
גבולות הסנכרון התומך וסינון מוגבל
On the Limits of Support-Preserving Alignment and Bounded Filtering
חוקרים בדקו את היכולת של מודלים גדולים לשפה, כגון LLM, לחדור ולהפיץ התנהגות מזיקה. המחקר מראה כי סינון מוגבל אינו מסוגל להסיר לחלוטין התנהגות מזיקה, אפילו עם עיבוד נוסף.
תקציר מקורי באנגליתarXiv:2607.18295v1 Announce Type: new Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research suggests that harmful behaviors can persist under preference-based alignment and that external filtering can be computationally hard in the worst case, but it remains unclear whether practical alignment pipelines that largely preserve internal representations can eliminate harmful behavior entirely rather than merely suppressing its most visible forms. We formalize this setting using support-preserving alignment operators together with bounded filtering algorithms under black-box, white-box, and statistical-query access,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית