כתבה
arXiv cs.CL ·
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
תקציר מקורי באנגליתarXiv:2609.09113v2 Announce Type: replace-cross Abstract: While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית