כתבה
arXiv cs.CL ·
אימון חופשי לפיענוח ספקולטיבי
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
שיטה חדשה לפיענוח ספקולטיבי ללא אימון, AdaptiveSpec, משפרת ביצועים על פני שיטות קודמות. השיטה מומלצת על ידי SGLang ומשתמשת במודלים כמו DeepSeek-R1-Distill-Llama-8B ו-Qwen3-8B.
תקציר מקורי באנגליתarXiv:2609.02897v2 Announce Type: replace Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית