כתבה
arXiv cs.AI ·
האנטומיה הפנימית של בחירה סטרטגית במודלי שפה גדולים
The Internal Anatomy of Strategic Choice in Large Language Models
מודלי שפה גדולים פועלים כאגנטים סטרטגיים ודגמים של בחירה אנושית. המחקר חקר את האנטומיה הפנימית של בחירה סטרטגית במודלי שפה גדולים.
תקציר מקורי באנגליתarXiv:2609.07478v1 Announce Type: new Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית