כתבה
arXiv cs.CL ·
הפענוח של המבקר: אופטימיזציה פרסונלית חופשית ללא פרסומים
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
אופטימיזציה פרסונלית חופשית ללא פרסומים למודלי LLM לאחר הכשרה. פיתוח טכנולוגיות חדשות לאופטימיזציה של מודלי LLM.
תקציר מקורי באנגליתarXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later tra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית