כתבה
arXiv cs.AI ·
Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates
תקציר מקורי באנגליתarXiv:2602.10430v2 Announce Type: replace-cross Abstract: Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures. We show that repeated reuse can turn this useful signal into excessive repulsion: as the learner moves away from a historical negative action, subsequent updates make that action increasingly remote without necessarily reducing its update strength. Our aggregate theory characterizes the resulting transition from a stable displacement beyond the positive-only target to persistent drift and the loss of finite stable equilibria; controlled strength sweeps show that an intermediate displacement can improve held-out reward. The relevant learner-relative coordinate is squared standardized distance for Gaussian poli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית