כתבה
arXiv cs.LG ·
המרה של מודל AR לדיפוזיה שומרת על יכולת יצירה
Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs
חוקרים מצאו דרך להמיר מודלים אוטורגרסיביים (AR) למודלי דיפוזיה, תוך שמירה על יכולת היצירה. השיטה מאפשרת יצירה מקבילה ללא צורך באימון מחדש. הניסויים הראו שהמודל החדש שומר על 95% מיכולת ה-GSM8K ו-99% מיכולת ה-MMLU-Pro של המודל המקורי.
תקציר מקורי באנגליתarXiv:2610.02657v1 Announce Type: new Abstract: Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the froz
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית