כתבה
arXiv cs.LG ·
זיהוי סיבתי-מודע של תיקון זיכרון פוקטואלי במודלי שפה MoE דלי
Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models
במאמר זה, חוקרים חוקרים את תהליך הזיכרון הפוקטואלי במודלי שפה MoE דלי. הם חוקרים את השאלה האם תיקון של זיכרון פוקטואלי יכול להיות מקומי למומחה אחד או תלוי בקבוצת המומחים שנרוטלים.
תקציר מקורי באנגליתarXiv:2606.03780v2 Announce Type: replace-cross Abstract: Activation patching can identify a mixture-of-experts (MoE) block whose clean output restores a corrupted factual prediction. However, because the block output combines contributions from multiple routed experts, block-level rescue does not establish whether the recovery localizes to an individual expert or depends on the routed expert set. We study this question on single-token COUNTERFACT contrasts by corrupting subject-token embeddings, restoring clean block outputs, and then restoring clean-minus-noised expert updates under fixed routing. In Qwen3-30B-A3B-Base, a discovery sweep selects layer 44, and held-out analysis identifies L44E069 as a recurrent routed contributor with positive specificity over same-layer active experts. I
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית