כתבה
arXiv cs.LG ·
הצפנה רציפה של גרדיאנט פרויקטיבי ללמידת הדמיה בטוחה
Interleaved Projected Gradient Descent for Safe Imitation Learning
למידת הדמיה תחת תנאי בטיחות. המאמר מציג תכנית ללמידת הדמיה של פוליציות נוירליות תחת תנאי בטיחות. התכנית משלבת צעדי גרדיאנט רגולרים עם צעדי בטיחות. המאמר מדגים את התכנית בתרגיל רכיבה אוטונומי.
תקציר מקורי באנגליתarXiv:2610.07167v1 Announce Type: cross Abstract: We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of $k$ safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting $k$ grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית