יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כמה דברים אומרים המעגלים? הערכת יציבות ומיוחדות של מעגלי ה-LM

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
במאמר זה, חוקרים חוקרים את יציבות ומיוחדות של מעגלי ה-LM. הם מצאו שמעגלים ברמת השבבים הם יציבים, אך לא מיוחדים, ומעגלים ברמת הנוירונים הם מיוחדים, אך לא יציבים.
תקציר מקורי באנגליתarXiv:2605.08348v3 Announce Type: replace Abstract: The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important for most tasks, but they are not specific: for a given task, ablating its own circuit drops accu
קרא במקור המקורי