יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Diffuzia פחות תמידית: פרקטיקה חדשה לדיפוזיה

Less Uniform Discrete Diffusion is More Powerful and Scalable
Diffuzia פחות תמידית היא פרקטיקה חדשה לדיפוזיה שמשפרת את הגמישות ואת תפוקת הדורש. היא כוללת חידושים באביזרי האימון ובאביזרי הדיפוזיה.
תקציר מקורי באנגליתarXiv:2609.35817v1 Announce Type: cross Abstract: Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B,
קרא במקור המקורי