יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התשובה הנכונה, תוצאה שגויה: זיהוי פריחה סטריאטית בהתערבויות סיבתיות

Right Answer, Wrong Mechanism: Detecting Pernicious Divergence in Causal Interventions
במאמר זה, החוקרים חקרו את השפעת התערבויות סיבתיות על רשתות עצביות. הם גילו שאלו התערבויות עשויות לגרום לתוצאות שגויות, ופיתחו תכונה חדשה שמסוגלת לזהות את התערבויות אלה.
תקציר מקורי באנגליתarXiv:2609.39243v1 Announce Type: new Abstract: Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expected answer through the wrong mechanism. No method currently tells the two cases apart. We make this question testable by planting hidden pathways inside pretrained language models; the pathways are silent on every benchmark prompt by construction, so which interventions depend on them is known exactl
קרא במקור המקורי