כתבה
arXiv cs.LG ·
Consensus Group Relative Policy Optimization for Text Generation
תקציר מקורי באנגליתarXiv:2602.03102v2 Announce Type: replace Abstract: Many strong decoding methods for text generation follow a sample-and-rerank paradigm: they draw multiple candidates, score each under a utility (reward) function using consensus across samples, and return the best one. Although effective, these methods incur high computational costs during inference due to repeated sampling and scoring. Prior attempts to amortize inference-time computation typically rely on gold references, teacher labels, or curated preference data, increasing dataset construction effort and the demand for high-fidelity reward models. We propose Consensus Group Relative Policy Optimization (C-GRPO), which distills Minimum Bayes Risk (MBR) decoding into training by formulating the consensus utility as a group-relative obj
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית