כתבה
arXiv cs.CL ·
Osprey: שיפור דראפטרים באמצעות פריטריינינג תלוי-מטרה
Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
Osprey משפר את ביצועי הדראפטרים באמצעות פריטריינינג תלוי-מטרה. המודל מאפשר הכללה רחבה יותר ושיפור ביצועים עבור מודלים כמו Llama ו-Qwen. הפריטריינינג מתבצע על בסיס מודל שפה קטן, ולאחר מכן מותאם לכל מודל יעד בנפרד.
תקציר מקורי באנגליתarXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which instead bootstraps drafters from off-the-s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית