יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בקרת האנטרופיה בתהליך התכלס

Divergence controls entropy in distillation
חוקרים בדקו את התהליך של התכלס במודלים גדולים של שפה, ומצאו כי האנטרופיה של המודל הסטודנט תלויה בנתונים ובפער שבין המודל המורה למודל הסטודנט. המחקר מראה כי הפער הזה יכול לשמש כרגולטור מפורש של האנטרופיה.
תקציר מקורי באנגליתarXiv:2610.03529v1 Announce Type: new Abstract: Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lo
קרא במקור המקורי