כתבה
arXiv cs.CL ·
Strangers to Themselves: What Language Models Say About Themselves Is Generic
תקציר מקורי באנגליתarXiv:2609.09899v1 Announce Type: cross Abstract: Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית