כתבה
arXiv cs.LG ·
Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
תקציר מקורי באנגליתarXiv:2610.10411v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית