יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת רישום רב-מוטיבציונית עמידה לסיכון מורכבת

Robust Risk-Sensitive Reinforcement Learning from Corrupted Human Feedback
למידת רישום רב-מוטיבציונית עמידה לסיכון מורכבת, המתמודדת עם פריפריה קורובטית של פריפרנסיות אנושיות.
תקציר מקורי באנגליתarXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under adversarial preference-label flips. We consider additive linear rewards and a fixed-reference protocol with one comparison per episode and at most $C$ flipped labels over $K$ episodes. We propose weighted streamed-preference CVaR RLHF (WSP-CVaR-RLHF), which combines uncertainty-weighted reward estimation with optimistic augmented-state CVaR planning. For known transitions and normalized rewards, we establish the regret bound $\widetilde{O}\left(\frac{d}{\kappa}\sqrt{\frac{K}{\alpha}}+\frac{dC}{\kappa\alpha}\right)$ u
קרא במקור המקורי