יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעבר לקיצור אוקלידיאני: התגברות על קריסת התגלמות ב-RL של LLM דרך RIPO

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
במאמר זה, המחברים מציגים פתרון לבעיה של קריסת התגלמות ב-RL של LLM. הם מציגים את RIPO, שהוא אלגוריתם של RIPO שמבטיח שגיאות פוליציה יסומנו באופן איזומטרי על המניפולד Riemannian.
תקציר מקורי באנגליתarXiv:2607.10169v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fundamental flaw of PPO-Clip: it implicitly measures policy discrepancy using Euclidean metric, which is theoretically inconsistent with the intrinsic geometry on the policy Riemannian manifold. This geometric mismatch results in overly conservative updates in low-probability regions while aggressive in high-probability regions, ultimately collapsing exploration. To correct this geometric flaw, we propose Riemannian Isom
קרא במקור המקורי