יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עיצוב פעולות: מדיניות סופגת מה שהיא יכולה לבטא

Action Shaping: Policies Absorb What They Can Express
חוקרים גילו עקרון חדש לעיצוב פעולות, שבו מדיניות סופגת אופסט שהיא יכולה לבטא בדיוק. המחקר מציג תנאי לספיגה, אלגוריתם ליישום ואבחון לקביעת עלות הסרת האופסט.
תקציר מקורי באנגליתarXiv:2609.32752v2 Announce Type: replace Abstract: Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then
קרא במקור המקורי