כתבה
arXiv cs.AI ·
ChronoSRL: גאומטריה זמנית ללמידת רפלקסיה עצמית
ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning
ChronoSRL היא פרדיגמה חדשה ללמידת רפלקסיה עצמית שמעניקה גאומטריה זמנית להתאמת המבקר. היא מתמקדת בעיצוב נתוני המבקר כדי לשפר את הביצועים של המדלג.
תקציר מקורי באנגליתarXiv:2609.36238v1 Announce Type: new Abstract: A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית