כתבה
arXiv cs.LG ·
Fully Online Decentralized Learning in Stochastic Games with Unknown Independent Chains
תקציר מקורי באנגליתarXiv:2610.01181v1 Announce Type: new Abstract: We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs. We develop a fully online, decentralized, and uncoordinated mirror-descent algorithm that operates in the dual space of occupancy measures for approximating stationary Nash equilibrium (NE) policies. The algorithm uses a single transition/reward sample at every primitive time step, relies only on local information, and requires neither coverage of the joint state space nor synchronized episodes. Under uniform-ergodicity and finite-coverage assumptions, we show that, with high probability, the time-averaged fixed-comparator regret decays at the canonical $O(T^{-1/2})$ rate, up to lo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית