כתבה
arXiv cs.AI ·
Carryover Drafting: Recycling Rejected States for Speculative Decoding
תקציר מקורי באנגליתarXiv:2609.14717v1 Announce Type: cross Abstract: Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We find that these discarded hidden states generated during target forward retain useful information about future tokens that can improve subsequent drafts. However, realizing this opportunity poses two distinct challenges. At inference, recycling overhead can increase drafting latency, diminishing the speedup gained from in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית