יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

העוני של פירושיות מכניסטיות: נקודת מבט רשמית

The Misery of Mechanistic Interpretability: A Formal Perspective
מספרים חדשים לבדיקת פירושיות של מודלי שפה. פיתוח רשמי של כלי לבדיקת פירושיות של מודלי שפה. פיתוח רשמי של כלי לבדיקת פירושיות של מודלי שפה.
תקציר מקורי באנגליתarXiv:2609.15533v1 Announce Type: new Abstract: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes. To gain insights into these models, interpretable replacement networks (IRNs) are trained at all layers, exposing interpretable features through sparsely activated neurons. However, the faithfulness of an IRN is usually evaluated only empirically on clean data, and we show that even semantically minor input perturbations flip the dominant IRN features-and thus the human-understandable interpretation-across five open-weight model families (GPT-2 small, Gemma 2 2B, Gemma 3 1B, Llama 3.2 1B, R1-Distill-Qwen 1.5B). We propose the first formal verification framework for the faithfulne
קרא במקור המקורי