כתבה
arXiv cs.LG ·
כאשר ניתן להסברה חיצונית לעזור ל-LLM להסביר?
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
במה דעמת סיוע חיצוני עוזר ל-LLM להסביר? חידוש חדש ב-RLLVR, GA-GRPO, גורס כי סיוע חיצוני יכול לשפר את ההסברה של LLM, אך יש לבחור את הכמות האופטימלית שלו. המאמר עוסק ב-GA-GRPO, פרקטיקה חדשה של RLLVR, שמשלבת סיוע חיצוני כדי לשפר את ההסברה של LLM.
תקציר מקורי באנגליתarXiv:2610.06861v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance} - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal. We close this gap with \emph{Guidance-Augmented GRPO} (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית