כתבה
arXiv cs.LG ·
איך הפרדה הופכת להחלטות הפוכות במודלי שפה מדוחסים
How Divergence Becomes Decision Flips in Compressed Language Models
מחקר חדש משווה בין הפרדה של KL לשונות כוללת במודלי שפה מדוחסים, ומוצא כי השונות הכוללת מנבאת יותר טוב את שיעור השינויים בהחלטות. המחקר בדק 19 מודלים פתוחים ו-9 משפחות הפרעה שונות.
תקציר מקורי באנגליתarXiv:2610.00694v1 Announce Type: cross Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed copies of 19 open models on five corpora and nine mechanically unrelated perturbation families, the rate at which the arg-max token changes (the \emph{flip rate}) tracks total variation at a ratio with median $1.05$, with no fitted constant. KL converts into flips only through its square root and a factor that varies fourfold across models and corpora, because KL averages over tokens before the root is taken; first-order stati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית