יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CAST: עקות יוצאות-על-עלות מעץ ספקולטיבי מדרגי-בלוק

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters
CAST: עקות יוצאות-על-עלות מעץ ספקולטיבי מדרגי-בלוק. זה עוזר למודלי השפה הגדולים להקדים טוקן עתידי זול ולאשר אותו עם המודל המטרה במקביל.
תקציר מקורי באנגליתarXiv:2610.00321v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs t
קרא במקור המקורי