כתבה
arXiv cs.CL ·
DLoop: Looped Speculative Decoding
תקציר מקורי באנגליתarXiv:2610.07659v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית