יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אימון דגמי סקיצה תפקודיים בצורה ישירה על ידי הקטנת התחושות הממוצעות של סיבובי פיזור

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
במאמר זה, המחברים מציגים פרקטיקה חדשה לאימון דגמי סקיצה תפקודיים. הם מציעים לאמן את הדגמים בצורה ישירה על ידי הקטנת התחושות הממוצעות של סיבובי פיזור. הם מציגים תאוריה תאורטית לאימון וביקורת של דגמי סקיצה תפקודיים על ידי ייצוג פיזור תפקודי כתהליך שכר תורתי. זה נותן כלי חדש לאימון דגמי סקיצה תפקודיים, והם מציגים תוצאות מעודכנות של דגמי סקיצה תפקודיים שאומנו באמצעות כלי חדש זה.
תקציר מקורי באנגליתarXiv:2610.10411v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for trainin
קרא במקור המקורי