כתבה
arXiv cs.CL ·
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
תקציר מקורי באנגליתarXiv:2605.10971v2 Announce Type: replace-cross Abstract: Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation methods for DLMs, imported from autoregressive models, apply uniform intervention at every denoising step. We show this uniform schedule is inefficient and degrades quality, and the damage compounds when multiple attributes are steered jointly. To diagnose the failure, we train sparse autoencoders on four DLMs (124M-8B parameters) and find that different attributes commit on distinct schedules, varying in timing, sharpness, and magnitude. For instance, topic commits within the first 2% of denoising on MDLM, whereas sentiment emerges gradually over 20% of t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית