יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

העברת רצף מחשבה לתוך מודלי שפה

Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models
חוקרים הציגו שיטה חדשה להפחתת זמן תגובה של מודלי שפה, באמצעות העברת רצף מחשבה לתוך המודל. השיטה, הנקראת 'העברה עצמית מסומנת', מאפשרת למודלים לייצר תשובות מדויקות יותר, תוך הפחתת הזמן והזיכרון הנדרשים. החוקרים בדקו את השיטה על מודל Qwen וגילו שהיא משפרת את ביצועי המודל.
תקציר מקורי באנגליתarXiv:2607.22629v3 Announce Type: replace Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises an obvious question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers with much shorter intermediate traces? We propose masked self-distillation, a knowledge-distillation based post-training framework in which copies of the same model are instantiated as teacher and studen
קרא במקור המקורי