כתבה
arXiv cs.CL ·
FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
תקציר מקורי באנגליתarXiv:2609.36202v1 Announce Type: new Abstract: Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית