כתבה
arXiv cs.LG ·
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
תקציר מקורי באנגליתarXiv:2609.36750v1 Announce Type: new Abstract: Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and fou
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית