כתבה
arXiv cs.LG ·
Co-Evolving LLM Evaluators and Policies via DynamicRubric
תקציר מקורי באנגליתarXiv:2607.20083v1 Announce Type: new Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a bottleneck for policy optimization: collapsed relative evaluator score gaps yield weak or misleading policy supervision. We theoretically characterize why these gaps matter through a probability allocation view, showing that the directional gain of shifting probability mass from one response to another is exactly the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. Motivated by this view, we propose DynamicRubric, a response-set-conditioned
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית