כתבה
arXiv cs.AI ·
על חלוקת נורמליזציה קדם-קדם בלמידת כוח
On BatchNorm Forward Modes in Value-Based Reinforcement Learning
במאמר זה, נראה כיצד חלוקת נורמליזציה קדם-קדם (BatchNorm) יכולה לשפר את רמת הביצועים בלמידת כוח ערך-ביסומטרי.
תקציר מקורי באנגליתarXiv:2609.06421v1 Announce Type: cross Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation. In C51, switching the BN bootstrap forward to batch-statistic mode significantly improves performance over unnormalized and LayerNorm baselines and scales stably with update-to-data ratios up to 12. In PQN, using batch-statistics for both ac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית