כתבה
arXiv cs.CL ·
Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
תקציר מקורי באנגליתarXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית