כתבה
arXiv cs.LG ·
Block Optimism for Nonstationary Bandits with Latent Linear Dynamics
תקציר מקורי באנגליתarXiv:2610.00911v1 Announce Type: cross Abstract: We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves $\tilde{O}(T^{2/3})$ regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית