כתבה
arXiv cs.LG ·
ReST-RL: שיפור תפיסה וסיכויים ב-LLMs
ReST-RL: Reinforcing LLM Reasoning through Unified Self-Training and Value-Guided Search
ReST-RL: שיפור תפיסה וסיכויים ב-LLMs. המאמר מציג פרקטיקה חדשה לשיפור תפיסה וסיכויים ב-LLMs, כולל שימוש ב-GPT-5 ו-Gemini.
תקציר מקורי באנגליתarXiv:2508.19576v3 Announce Type: replace-cross Abstract: With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method - Group Relative Policy Optimization (GRPO) - has achieved critical success, yet it still suffers from the issue of insignificant reward signals. This paper introduces ReST-RL, a unified Reinforced Self-Training (ReST) policy-value framework that reconnects policy optimization and value-guided search to improve LLM reasoning ability. Firstly, ReST-GRPO adopts an optimized ReST-style algorithm to reshape the policy-induced trajectory distribution by increasing the reward variance of GRPO sampling and exposing the policy to more informative partial states, thereby improving training efficiency and effectiveness. Then, we fur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית