כתבה
arXiv cs.CL ·
כאשר חידוש חיציוני עוזר לתפיסה של LLM?
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
במה דעת חיציוני חיציוני עוזר לתפיסה של LLM? המאמר עוסק בשאלה זו ומציג תאוריה חדשה של חידוש-מסגרת GRPO.
תקציר מקורי באנגליתarXiv:2610.06861v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance} - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal. We close this gap with \emph{Guidance-Augmented GRPO} (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית