יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Pivot-SD: פיתוח עצמי יעיל להתמסרות מסוגלת לדיפוזיה של מודלי שפה

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD הוא פיתוח עצמי יעיל שמשמש להתמסרות מסוגלת לדיפוזיה של מודלי שפה. הוא נועד לטפל בבעיה של קביעת קרדיט, כאשר כמה התחייבויות במהלך תהליך התיקון קובעות את רוב התשובה. Pivot-SD נבחן באמצעות ניסויים שבהם הוא השתמש במודל LLaDA-8B-Instruct והציג תוצאות טובות יותר מאשר פסקאות תרגול שלמות.
תקציר מקורי באנגליתarXiv:2610.03665v1 Announce Type: cross Abstract: Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an inform
קרא במקור המקורי