כתבה
arXiv cs.LG ·
OPIUM: הפחתת השפעות חיצוניות וסירוב יתר
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
OPIUM היא שיטה לטיהור וקטורי הכוונה במודלי שפה גדולים. היא משפרת את הבטיחות והיעילות של המודלים בזמן פעולה.
תקציר מקורי באנגליתarXiv:2607.19806v1 Announce Type: new Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts. We introduce OPIUM (Optimizing Protected Injections via Utility Manifolds), a training-free method for sanitizing steering vectors through representation matching. Given reference behaviors on two prompt sets, OPIUM optimizes a new steering vector that preserves the downstream representations induced by the desired intervention while matching a safer reference behavior on prompts where the original vector fails. Across steering-externality and over-refusal settings, OPIUM
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית