כתבה
arXiv cs.LG ·
מתי פעולות-תשומת-לב חיסול תומכות טענות קאוזליות? קונפונדים-הצפנה, רצפות-קרקע, ובקרות-מתאימות
When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls
פעולות-תשומת-לב חיסול עשויות להיות רגישות ללא תכנון-מערכת נכון. נמצא כי זיהוי-ראש-העין-חיסול אינו גורם לטענות-קאוזליות על ידי עצמו, ודרושה תכנון-מערכת נכון, מדד-ביצועים-לא-מצומצם, ובקרות-מתאימות.
תקציר מקורי באנגליתarXiv:2610.00373v1 Announce Type: new Abstract: Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,00
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית