כתבה
arXiv cs.LG ·
AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
תקציר מקורי באנגליתarXiv:2607.19223v1 Announce Type: new Abstract: Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central pitfall of diffusion drafters: bidirectional attention is a double-edged sword. On one hand, it endows the model with parallel generation and global contextual modeling capabilities; on the other hand, this inherent global dependency introduces high variance at both the domain-level and the token-level: acce
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית