כתבה
arXiv cs.LG ·
ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
תקציר מקורי באנגליתarXiv:2607.03126v3 Announce Type: replace Abstract: Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their unequal contributions to the reasoning process. Entropy provides a natural indicator of the model's decision state, yet using it for token-level credit assignment presents two key challenges: long-tail probabilities in large vocabularies corrupt both entropy values and gradients, and uncertainty carries distinct semantics across positive- and non-positive-advantage trajectories. We propose Asymmetric Credit Policy Optimization (ACPO), which replaces global entropy with the complement of the top-token probability as a mode-local proxy. Guided by gradient analysi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית