יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הפצת זמן: דיסטילציה עצמית למודלי שפה דיפוזיונליים

Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
מודלי שפה דיפוזיונליים: הפצת זמן ודיסטילציה עצמית למודלי שפה דיפוזיונליים. המאמר עוסק בפיתוח של טכניקה חדשה להפחתת זמן עבודה של מודלי שפה דיפוזיונליים. הטכניקה, הקרויה 'הפצת זמן', מאפשרת למודלי שפה דיפוזיונליים להפחית את זמן העבודה שלהם באופן ניכר. המאמר כולל תיאור של הטכניקה, תיאור של התוצאות המעשיות שלה, ודיווח על ניסויים שנערכו.
תקציר מקורי באנגליתarXiv:2609.15177v1 Announce Type: new Abstract: Diffusion language models (dLLMs) promise fast inference by generating multiple tokens in parallel, but suffer severe performance degradation when parallel decoding is pushed too aggressively. We introduce Temporal Self-Distillation (TSD), a simple on-policy method that trains dLLMs for fast inference by distilling predictions across time. Specifically, TSD distills the model's denoising distribution at earlier timesteps toward its distribution at the final timestep at which a token is committed. This encourages earlier predictions to better anticipate the model's eventual output, enabling much more aggressive parallel decoding. Because its teacher signal comes from the model itself, TSD requires no offline teacher generation and applies seam
קרא במקור המקורי