כתבה
arXiv cs.AI ·
צורה טובה יותר להתחלה מכל יעד
A Better Spur Should Start From Each Objective
MMPO מציעה פתרון לבעיות התפתחותיות ב-MORL, על ידי פיצול גרדיאנטים והגבלות עצמיות. התוצאות נראות טובות בנתחים ריאל-עולם.
תקציר מקורי באנגליתarXiv:2609.08211v1 Announce Type: new Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing traditional linear scalarization to experience severe metric oscillations. To address optimization conflicts among multiple objectives in real-world deployment scenarios, we propose Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization. Specifically, MMPO performs exposure debiasing to mitigate sparse and biased rewards, applies priority-aware orthogonal projection to decouple conflicting gradients, and introduces self-prompted gradient constraints to preven
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית