יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Pivot-SD: עיבוי עצמי יעיל למודלים שפה דיפוזיים

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD הוא כלי עיבוי עצמי יעיל למודלים שפה דיפוזיים. הוא משפר את LLaDA-8B-Instruct על ידי אימון רק על החלטות בעלות השפעה גבוהה. השיפור נמדד ב-200 שאלות ו-4 רולאוטים, והוא עולה על שיטות אימון אחרות.
תקציר מקורי באנגליתarXiv:2610.03665v1 Announce Type: new Abstract: Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an informat
קרא במקור המקורי