יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

WASD: העברת ידע מבוססת Wasserstein למודלי שפה גדולים

WASD: Wasserstein-based Knowledge Distillation for Large Language Models
WASD היא שיטה חדשה להעברת ידע ממודלי שפה גדולים למודלים קטנים יותר. השיטה משתמשת במרחק Wasserstein ובמטריצת עלות המבוססת על הטמעות טוקנים. הניסויים הראו ש-WASD משפר את ביצועי ההעברה במשימות שונות, כולל הוראות, נימוק מתמטי ויצירת קוד.
תקציר מקורי באנגליתarXiv:2610.07706v1 Announce Type: cross Abstract: Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost
קרא במקור המקורי