כתבה
arXiv cs.AI ·
Local Sparsity Enables Unsupervised LLM Safety Detection
תקציר מקורי באנגליתarXiv:2609.20129v2 Announce Type: replace-cross Abstract: Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE)
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית