כתבה
arXiv cs.AI ·
אמון, אל תאמין, או פליפ: למידת תגמול ברובוסטי בהתבסס על ניסיון
Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback
למידת תגמול ברובוסטי המשתמשת בפידבק משותף ממומנן
תקציר מקורי באנגליתarXiv:2601.18751v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PBRL) offers a promising alternative to explicit reward engineering by learning from pairwise trajectory comparisons. However, real-world preference data often comes from heterogeneous annotators with varying reliability; some accurate, some noisy, and some systematically adversarial. Existing PBRL methods either treat all feedback equally or attempt to filter out unreliable sources, but both approaches fail when faced with adversarial annotators who systematically provide incorrect preferences. We introduce TriTrust-PBRL (TTP), a unified framework that jointly learns a shared reward model and expert-specific trust parameters from multi-expert preference feedback. The key insight is that trus
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית