כתבה
arXiv cs.LG ·
מודל השפה שלו הוא ביקורתו: אימון רפלקסיבי עם הערכת ערך ממצבי המערכת
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
מודל השפה שלו משמש כביקורת עצמית לאימון רפלקסיבי. המחקר מציג חידוש באימון רפלקסיבי עם הערכת ערך ממצבי המערכת. המודל השפה 'Gemini' משמש כדוגמה למודל השפה שלו. המחקר נעשה בשיתוף עם 'GPT-5'.
תקציר מקורי באנגליתarXiv:2605.07579v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimation),
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית