כתבה
arXiv cs.LG ·
Reward Valuation in Large Language Models: Causal Induction of Anhedonia
תקציר מקורי באנגליתarXiv:2607.06626v2 Announce Type: replace Abstract: Recent frontier models mimic complex aspects of human cognition. Here we ask whether this alignment extends into reward valuation, which we assess in a mechanistic framework. Specifically, we use clinical tests that were developed to evaluate anhedonia in human subjects with major depressive disorders. Mechanistically, anhedonia is frequently associated with dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in state-of-the-art AI models, and evaluate their causal in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית