כתבה
arXiv cs.CL ·
זיהוי סיבתי חכם של זיכרון פקטואלי במודלי שפה MoE דלילים
Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models
במאמר זה, החוקרים חקרו את השאלה האם תיקון חלק של מודל MoE, שמשחזר פלט חסר תקלה, נובע מאחד המומחים המשתתפים או מהקבוצה של מומחים שנשלחו. התוצאות הראו שתיקון החלק לא תמיד מעיד על התמקדות במומחה יחיד.
תקציר מקורי באנגליתarXiv:2606.03780v2 Announce Type: replace Abstract: Activation patching can identify a mixture-of-experts (MoE) block whose clean output restores a corrupted factual prediction. However, because the block output combines contributions from multiple routed experts, block-level rescue does not establish whether the recovery localizes to an individual expert or depends on the routed expert set. We study this question on single-token COUNTERFACT contrasts by corrupting subject-token embeddings, restoring clean block outputs, and then restoring clean-minus-noised expert updates under fixed routing. In Qwen3-30B-A3B-Base, a discovery sweep selects layer 44, and held-out analysis identifies L44E069 as a recurrent routed contributor with positive specificity over same-layer active experts. Its eff
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית