כתבה
arXiv cs.LG ·
Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data
תקציר מקורי באנגליתarXiv:2609.38880v1 Announce Type: cross Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $\gamma$, write $H=(1-\gamma)^{-1}$, and let $\mu_{\min}$ and $t_{\operatorname{mix}}$ denote the minimum stationary probability and total-variation mixing time. We prove that last-iterate TD achieves sup-norm error at most $\varepsilon$ with high probability using$\widetilde O\left( \frac{H^3}{\mu_{\min}\varepsilon^2} +\frac{t_{\operatorname{mix}}}{\mu_{\min}} \right)$ transitions, for $0<\varepsilon\leq1$. This rate holds both for a constant step size selected for the target accuracy and for a decreasing schedule independent of the target accuracy and terminal time. The latter give
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית