יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חיזוק שרשרת המחשבה: הפקת LLM עם אילוצי תגמול

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
חוקרים פיתחו שיטה חדשה להפקת מודלים LLM קטנים יותר עם יכולות תיאור טובות יותר. השיטה משתמשת בלמידת תגמול מוגבלת כדי לחזק את השרשרת המחשבתית של המודל. התוצאות מראות שיפור משמעותי בדיוק וביכולת המודל לייצר קוד.
תקציר מקורי באנגליתarXiv:2610.00332v1 Announce Type: cross Abstract: Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest
קרא במקור המקורי