כתבה
arXiv cs.CL ·
EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation
תקציר מקורי באנגליתarXiv:2605.29847v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has advanced Large Language Models (LLMs) in verifiable domains, while open-ended generation remains challenging due to the absence of definitive rewards. Rubric-based RL provides explicit evaluation criteria, but learning to construct these criteria remains challenging when final-answer correctness is not verifiable. We propose EvoRubric, a co-evolutionary RL framework that combines criterion-validity feedback, response discrimination, and peer agreement to learn rubrics for open-ended generation. A shared policy acts as both a Reasoner and a Rubric Generator, using its current responses and historical rubrics to discover new evaluation dimensions. To combine adaptive rubric discovery with a stable validity ch
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית