יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

DAMPER: ניהול גרדיאנטים לשם פוליציות חלקות

DAMPER: Return-Prioritized Gradient Control for Smooth Policies
DAMPER הוא שיטה חדשה לניהול גרדיאנטים שמטרתה לייצר פוליציות חלקות. השיטה משלבת גרדיאנטים שונים ומשפרת את התפלגות הפוליציות. המחברים הציגו את DAMPER במאמרם והציעו דוגמאות לשימוש בשיטה.
תקציר מקורי באנגליתarXiv:2609.38903v1 Announce Type: new Abstract: Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible when their gradients are small relative to the native actor gradient. Moreover, existing methods often combine multiple auxiliary losses, complicating loss balancing without necessarily improving the return-smoothness trade-off. We introduce DAMPER (Direction-Aware Magnitude-Controlled Projection with Explicit Return Priority), which combines the native actor gradient with a temporal-consistency gradient through conflict-conditioned projection and adaptive magnitude control. It removes the auxiliary component opposi
קרא במקור המקורי