כתבה
arXiv cs.AI ·
PAIR: דגם פניקס-מודע לתגמול פנימי לאופטימיזציה של גורמים מרובע-פניה
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
דגם פניקס-מודע לתגמול פנימי לאופטימיזציה של גורמים מרובע-פניה. הדגם משתמש בפרפקס-אומדן כדי לספק תגמול פנימי לגורמים בסביבת פניה-מרובע. הדגם נבחן במספר סקנרים והוכח כמוצלח.
תקציר מקורי באנגליתarXiv:2605.17877v3 Announce Type: replace Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been emerging as a leading choice, but its reliance on sparse outcome rewards severely limits credit assignment across intermediate steps. Existing remedies such as running full rollouts to assign step-level advantages, calling external LLM judges at each step, or computing intrinsic rewards that require ground-truth answers at every evaluation introduce significant costs or practical constraints. We hypothesize that internal correctness probing over LLM hidden states can be repurposed as a step-level reward signal, potentially addressing all of these limitations at once. However, existing probing research as
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית