יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כלי פסיכומטרי ל-LLM

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
חוקרים פיתחו כלי פסיכומטרי ייחודי למודלים LLM. הם בדקו 25 מודלים ומצאו פער בין דיווח עצמי להתנהגות. המחקר מראה שהמודלים לא תמיד מדווחים באמת על התנהגותם.
תקציר מקורי באנגליתarXiv:2606.09843v4 Announce Type: replace-cross Abstract: Do large language models' (LLMs') answers to self-report questionnaires predict how they behave? Prior work finds they do not, but it uses human personality inventories, so the gap could reflect borrowed human constructs rather than LLM self-report itself. We test this with a self-report instrument built from LLM-specific behaviors (e.g., over-refusal, unsolicited disclaimers) whose structure is derived bottom-up. Administering 300 items 30 times to 25 LLMs from 17 developers yields five replicable, reliable factors (Tucker $\phi \geq .957$, $\alpha \geq .930$). We compare these self-reports with 2,500 open-ended behavioral samples rated by 151 humans and an LLM-judge ensemble. Humans and judges agree about model behavior ($\bar{r}
קרא במקור המקורי