כתבה
arXiv cs.LG ·
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
תקציר מקורי באנגליתarXiv:2609.30500v1 Announce Type: new Abstract: Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transforme
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית