כתבה
arXiv cs.LG ·
אופטימיזציה פארטו-אופטימלית ללמידת רפלקסיה מקוונית באופן רגולרי
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebycheff Scalarization
אופטימיזציה פארטו-אופטימלית ללמידת רפלקסיה מקוונית באופן רגולרי. המחקר מציג אלגוריתם חדש לאופטימיזציה של פריפרקציות פארטו-אופטימליות בלמידת רפלקסיה מקוונית.
תקציר מקורי באנגליתarXiv:2604.13175v2 Announce Type: replace Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objective alignment is well-studied, many real-world applications demand the simultaneous optimization of multiple conflicting rewards, e.g. activity and specificity for proteins, or helpfulness and harmlessness for chatbots. Prior work has largely relied on linear reward scalarization, which provably fails to recover non-convex regions of the Pareto front. In this paper, instead of scalarizing the rewards directly, we frame multi-objective RL itself as an optimization problem to be scalarized via smooth Tchebycheff scalarization, a recent technique that overcomes the drawbacks of linear scalariza
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית