כתבה
arXiv cs.LG ·
ניתוח דינמיקת יצירה של מודלי שפה דיפוזיה
Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
חוקרים פיתחו שיטה חדשה לניתוח דינמיקת יצירה של מודלי שפה דיפוזיה. השיטה, הנקראת Diffusion Layer Integrated Gradients, מאפשרת להבין כיצד מודלים אלה יוצרים טקסט. המחקר הראה כיצד השיטה יכולה לשמש לניתוח מודלים שונים.
תקציר מקורי באנגליתarXiv:2610.01177v1 Announce Type: cross Abstract: This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infill
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית