כתבה
arXiv cs.CL ·
In-Context Learning as Implicit Policy Gradient
תקציר מקורי באנגליתarXiv:2607.23153v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical foundations underlying this phenomenon remain poorly understood. In this paper, we show that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization. We first provide a constructive proof that self-attention mechanisms can implement reward-weighted aggregation analogous to the REINFORCE algorithm under specific weight matrix configurations, and discuss the relationship between this construction and the behavior of pretrained transformers. The cor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית