יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

איך התפלגות הופכת להחלפת החלטות במודלי שפה מדוחסים

How Divergence Becomes Decision Flips in Compressed Language Models
מחקר חדש משווה בין התפלגות KL לשונות כוללת במודלי שפה מדוחסים. התוצאות מראות כי שונות כוללת מנבאת יותר טוב את שיעור החלפת ההחלטות. הממצאים עשויים לשפר את הבנתנו במודלים מדוחסים.
תקציר מקורי באנגליתarXiv:2610.00694v1 Announce Type: cross Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed copies of 19 open models on five corpora and nine mechanically unrelated perturbation families, the rate at which the arg-max token changes (the \emph{flip rate}) tracks total variation at a ratio with median $1.05$, with no fitted constant. KL converts into flips only through its square root and a factor that varies fourfold across models and corpora, because KL averages over tokens before the root is taken; first-order stati
קרא במקור המקורי