כתבה
arXiv cs.AI ·
זרים לעצמם: מה שמודלי השפה אומרים על עצמם הוא כללי
Strangers to Themselves: What Language Models Say About Themselves Is Generic
מודלי השפה יכולים לתאר את עצמם באופן פורפורמצי, אך התיאורים אינם ספציפיים למודל. המחקר מצא כי התיאורים של המודלים על עצמם דומים לתיאורים של 'אגנטים AI יעילים' בכלל, וכי המודלים נוטים לתאר את עצמם באופן חיובי.
תקציר מקורי באנגליתarXiv:2609.09899v1 Announce Type: cross Abstract: Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית