יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

פיקוח פעולה מסדר גבוה יוצר מדיניות חזקה

Higher-Order Action Supervision Makes A Strong Policy Class
שיטה חדשה לפיקוח פעולה מסדר גבוה משפרת באופן משמעותי את ביצועי המדיניות בלמידת חיקוי ולמידת תגמול. השיטה מאפשרת פיקוח על פעולות מסדר ראשון ושני, מה שמשפר את יציבות הבקרה ואת העמידות ביישומים מעשיים.
תקציר מקורי באנגליתarXiv:2610.11175v1 Announce Type: cross Abstract: Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhanc
קרא במקור המקורי