יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת ריפוד עם פידבק של רגלי חתך תחת התאמה לינארית

Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation
למידת ריפוד עם פידבק של רגלי חתך תחת התאמה לינארית. המאמר חוקר את השפעת גודל החתך על הלמידה.
תקציר מקורי באנגליתarXiv:2610.08271v1 Announce Type: new Abstract: Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning. To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation. Our work answers how the granularity of segment feedback and the choice of segmentation influence learning. For equal-length segments with known transitions, we design algorithms $\bitssegd$ and $\edlinucbsegd$ for binary and sum feedback types, respectively. They
קרא במקור המקורי