יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קריקטורית חדשה ללמידת חיזוק

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning
חוקרים הציגו קריקטורית חדשה ללמידת חיזוק, BVER, המאפשרת לבוטים ללמוד מהר יותר. BVER מתרחבת משני הקצוות, מה שמאפשר לה לכסות פחות איטרציות. היא הראתה תוצאות טובות במשימות רובוטיות.
תקציר מקורי באנגליתarXiv:2610.03395v1 Announce Type: new Abstract: Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outwa
קרא במקור המקורי