כתבה
arXiv cs.LG ·
How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models
תקציר מקורי באנגליתarXiv:2609.34514v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framew
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית