כתבה
arXiv cs.LG ·
SEED: Self-Speculative Decoding via Implicit Encoder-Decoder
תקציר מקורי באנגליתarXiv:2609.36590v1 Announce Type: cross Abstract: Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed durin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית