יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תגמול חשיבה טובה יותר ל-LLM

Rewarding Better Thinking for LLM Preference Alignment
חוקרים מציגים שיטה חדשה לתגמול LLM, המכונה Thinking Checklist Reward, המיועדת לשפר את החשיבה וההתאמה להעדפות אנושיות. השיטה הוכחה כיעילה בניסויים.
תקציר מקורי באנגליתarXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response while providing limited guidance for the reasoning trajectory. This can make credit assignment coarse when multiple responses receive similar final scores, leaving trajectory-level preferences under-specified. To address this limitation, we propose Thinking Checklist Reward (TCR), a process-oriented reward for RL-based preference alignment. TCR converts preference pairs into sample-specific thinking checklists and uses them to evaluate whether the generated reasoning t
קרא במקור המקורי