יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת חיזוק קמור-קעור

Convex-Concave Reinforcement Learning
חוקרים הציגו שיטה חדשה ללמידת חיזוק, Convex-Concave RL, המאפשרת לעקוף את הבעיה הלא-קמורה של מקסימיזציה של תשואה צפויה. השיטה משתמשת בתכנות קמור-קעור ומאפשרת לפתור את הבעיה בצורה יעילה יותר.
תקציר מקורי באנגליתarXiv:2610.09108v1 Announce Type: new Abstract: Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[\pi/\pi_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This
קרא במקור המקורי