כתבה
arXiv cs.LG ·
קללת אופק בלמידת מדיניות
An Informational Curse of Horizon in Goal-Conditioned Policy Learning
חוקרים גילו קללת אופק חדשה בלמידת מדיניות מותנית ביעד, שמשפיעה על ביצועים וכלליות. הם מצאו שאימון על יעדים רחוקים יותר יכול לפגוע בביצועים, וששיטות למידת מדיניות שונות מושפעות אחרת מהתופעה.
תקציר מקורי באנגליתarXiv:2610.09247v1 Announce Type: new Abstract: The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performan
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית