יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Inducing Process Supervision from Outcome-Only Reinforcement Learning

תקציר מקורי באנגליתarXiv:2609.36641v1 Announce Type: new Abstract: Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relativ
קרא במקור המקורי