יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

גרדיאנטים יודעים מה תוצאות לא: הפתחה של למידת התנהגות לטווח ארוך ב-LLM עם פרסומים מאולצים

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
הצגנו פרסום מאולץ שמאפשר ל-LLM ללמוד לטווח ארוך עם פרסומים מאולצים. הפרסום נכתב בעזרת המודל Qwen.
תקציר מקורי באנגליתarXiv:2609.03342v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosin
קרא במקור המקורי