כתבה
arXiv cs.AI ·
מודל שפה כבודק: למידת חיזוק עם הערכת ערך ממצבים פנימיים
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
POISE הוא אלגוריתם למידת חיזוק המשתמש במצבים פנימיים של המודל להערכת ערך. הוא משפר את יציבות האימון ומגיע לתוצאות טובות יותר מאשר שיטות אחרות. POISE נבדק על מודלים Qwen3-4B ו-OLMo3-7B-Instruct-DPO.
תקציר מקורי באנגליתarXiv:2605.07579v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית