כתבה
arXiv cs.LG ·
מדוע ניתן לבנות פוליסית חזקה על נתונים רגילים?
From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
הצגנו פרוטוקול חדש ללמידה בעקיפה, המאפשר לבנות פוליסית חזקה על נתונים רגילים. הפרוטוקול, QTPT, נבחן במבחנים של RL והוכיח רובוסטנסיות חזקה יותר לאיכות הנתונים.
תקציר מקורי באנגליתarXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MD
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית