כתבה
arXiv cs.AI ·
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
תקציר מקורי באנגליתarXiv:2607.19806v1 Announce Type: cross Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts. We introduce OPIUM (Optimizing Protected Injections via Utility Manifolds), a training-free method for sanitizing steering vectors through representation matching. Given reference behaviors on two prompt sets, OPIUM optimizes a new steering vector that preserves the downstream representations induced by the desired intervention while matching a safer reference behavior on prompts where the original vector fails. Across steering-externality and over-refusal settings, OPIU
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית