כתבה
arXiv cs.AI ·
פשטות ייצוגית וגודל מעגל מתנתקים בדרך תלוית סף
Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
חוקרים בדקו אם פשטות ייצוגית ופירוק תכונות מרוכזות מנבאים גודל מעגל קטן יותר. הם השתמשו באימון עקבי ואימון עקבי עוין על מודל GPT-2 Small, ומצאו שהמודל העוין הוא יותר פשוט ודורש פחות קשתות ברמות אמינות גבוהות.
תקציר מקורי באנגליתarXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית