כתבה
arXiv cs.LG ·
On Trajectory-Aware Training for Masked Diffusion Language Models
תקציר מקורי באנגליתarXiv:2609.37974v1 Announce Type: new Abstract: Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית