כתבה
arXiv cs.AI ·
למידת חיזוק לא מקוונת עם אופטימיזציה רב-מטרות
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebycheff Scalarization
חוקרים פיתחו אלגוריתם חדש ללמידת חיזוק לא מקוונת עם אופטימיזציה רב-מטרות. האלגוריתם, STOMP, משתמש בסקלריזציה חלקה של Tchebycheff כדי לפתור בעיות רב-מטרות. הוא הוכח כיעיל במגוון ניסויים.
תקציר מקורי באנגליתarXiv:2604.13175v2 Announce Type: replace-cross Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objective alignment is well-studied, many real-world applications demand the simultaneous optimization of multiple conflicting rewards, e.g. activity and specificity for proteins, or helpfulness and harmlessness for chatbots. Prior work has largely relied on linear reward scalarization, which provably fails to recover non-convex regions of the Pareto front. In this paper, instead of scalarizing the rewards directly, we frame multi-objective RL itself as an optimization problem to be scalarized via smooth Tchebycheff scalarization, a recent technique that overcomes the drawbacks of linear sca
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית