כתבה
arXiv cs.LG ·
Robust Asynchronous Q-Learning under Reward and State Corruption via Batching
תקציר מקורי באנגליתarXiv:2607.20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observations of the learner following the Huber contamination model. To defend against such data corruption, we propose {{\texttt{BR-Async-Q}}}: a novel, epoch-based, robust \(Q\)-learning algorithm built upon two key ideas: (i) partitioning the online data stream into batches to reduce variance, and (ii) constructing robust estimates of the Bellman optimality operator using such batched data. We prove a high-probability $\ell_\infty$ error bound for {{\texttt{BR-Async-Q}}} that matches that for vanilla \
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית