כתבה
arXiv cs.LG ·
למידת תגמול מעדפות באמצעות מרחב תת-מרחב
Subspace Inference Enables Efficient Active Reward Learning from Preferences
חוקרים פיתחו שיטה חדשה ללמידת תגמול מעדפות אנושיות. השיטה, PreferenceEKF, משתמשת במרחב תת-מרחב כדי לעקוב אחר אי-ודאות במודל התגמול, מה שמאפשר למידה פעילה יעילה יותר. השיטה הוכחה כיעילה בניסויים על מאגרי הנתונים D4RL ו-V-D4RL.
תקציר מקורי באנגליתarXiv:2609.04066v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preference queries. However, effective uncertainty quantification required for active learning remains a key challenge for large neural network reward models. In this paper, we introduce PreferenceEKF, a sample-efficient approach that tracks reward model uncertainty by framing active preference learning as a sequential Bayesian filtering problem. Instead of relying on computationally prohibitive posterior inference over the full neural network parameter space, our method performs sequential inference via an extended
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית