כתבה
arXiv cs.LG ·
מעבר מ-Euclidean Clipping: פיצול התגלמות ב-LLM RL דרך Riemannian Isometric Policy Optimization
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
מאמר זה מציג פתרון חדש לבעיה של פיצול התגלמות ב-Learning Language Models (LLMs) דרך Riemannian Isometric Policy Optimization. הפתרון נבחן ב-7 מבחני תחרות והוכיח יתרון של 60% על GRPO ב-AIME24.
תקציר מקורי באנגליתarXiv:2607.10169v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fundamental flaw of PPO-Clip: it implicitly measures policy discrepancy using Euclidean metric, which is theoretically inconsistent with the intrinsic geometry on the policy Riemannian manifold. This geometric mismatch results in overly conservative updates in low-probability regions while aggressive in high-probability regions, ultimately collapsing exploration. To correct this geometric flaw, we propose Riemannian Isometric
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית