יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שחזור הנחיה מחוץ למדיניות לפיענוח ספקולטיבי

Recovering Off-Policy Supervision for Speculative Decoding
חוקרים הציגו שיטה חדשה לשחזור הנחיה מחוץ למדיניות לפיענוח ספקולטיבי. השיטה משתמשת במסגרת אימון מבוססת רולאוט כדי לשחזר הנחיה תקפה. התוצאות מראות שיפור משמעותי באורך הקבלה הממוצעת.
תקציר מקורי באנגליתarXiv:2609.38795v1 Announce Type: new Abstract: Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the
קרא במקור המקורי