כתבה
arXiv cs.LG ·
הקשר החלש: תיפוקטורינג LLM עם רגולציה חסרת-תועלת
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
במאמר זה, המחברים מציגים שיטה חדשה לתיפוקטורינג יכולות התפיסה של מודלי שפה גדולים (LLM) לתלמידים קטנים יותר. השיטה מבוססת על רגולציה חסרת-תועלת, שהיא רגולציה שמטרתה למנוע תועלת-הקשר, שהיא תופעה שבה התלמיד מגיע לתשובה נכונה על ידי סדרה של תהליכים לא נכונים. השיטה נועדה לשפר את יעילות השימוש ב-LLM ולאפשר תיפוקטורינג של יכולות התפיסה שלהם.
תקציר מקורי באנגליתarXiv:2610.00332v1 Announce Type: new Abstract: Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest li
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית