כתבה
arXiv cs.AI ·
SAEScientist-Bench: האם סוכני AI יכולים לבצע מחקר עצמאי ב-SAE?
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
סוכני AI יכולים לבצע מחקר עצמאי ב-SAE, אך נשארים מאחורי הבסיס המומחה. המחקר כולל ניתוח של SAEScientist-Bench, כלי לבדיקת יכולות סוכני AI ב-SAE. התוצאות הראו שהסוכנים יכולים לגלות תכונות ולעצב תרגילים, אך עדיין נשארים מאחורי המומחים בתחום. המחקר חשוב כיוון שהוא עוסק בפיתוח כלים לבדיקת יכולות סוכני AI ב-SAE, ומציע תוצאות שיכולות לשפר את ההבנה של סוכני AI.
תקציר מקורי באנגליתarXiv:2609.09113v1 Announce Type: new Abstract: While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית