כתבה
arXiv cs.AI ·
התקרבות גאומטרית ל-SAC עם זונוטופים ללמידת הליכה
A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning
במאמר זה, המחברים מציגים גישה גאומטרית חדשה ל-SAC (Soft Actor-Critic) שמשתמשת בזונוטופים ללמידת הליכה. הם מציעים את GeZo-SAC, שמשתמשת בייצוגים גאומטריים חיצוניים כדי להתאים את החשיבה הפוסטיביסטית של המבקרים למצב ולפעולה. הם מדגימים את GeZo-SAC בארבעה בסיסים של MuJoCo-v5 ללמידת הליכה ומציגים תוצאות טובות יותר מאשר שיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.12113v1 Announce Type: cross Abstract: Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית