כתבה
arXiv cs.LG ·
When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal
תקציר מקורי באנגליתarXiv:2607.11953v4 Announce Type: replace Abstract: Does a reinforcement-learning agent that earns reward learn its task's hidden state? We study this question with hidden finite automata that the agent partially controls. Because each automaton is known, we can normalize reward by the best achievable return and probe the network for the true state at every step. Together the two measurements separate failures that reward alone conflates. An agent can encode too little of a state its network could hold, or encode the state and still control poorly. Weak on-policy RL matches random play while the state probe stays at chance. State learning depends on the optimizer, the training budget, and the task's structure. Permutation automata provide a warning before training: no input symbol maps two
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית