יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כמה דברים יודעים המסלולים? - מדידת התאמה ומיוחדות של מסלולי השפה

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
במאמר זה, חוקרים חוקרים את התאמה ומיוחדות של מסלולי השפה במודלי LLM. הם מצאו שמסלולי השפה הם קבועים, אך לא מיוחדים.
תקציר מקורי באנגליתarXiv:2605.08348v2 Announce Type: replace Abstract: The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's per
קרא במקור המקורי