כתבה
arXiv cs.LG ·
אמונות והתנהגות במודלים של שפה
Beliefs and Behavior in Language Models
חוקרים בודקים אם מודלים גדולים של שפה (LLM) מחזיקים באמונות. הם מציעים שיטה לבדיקה אמפירית של שאלה זו, ומוצאים כי מודלים מתקדמים יכולים להיות מתוארים כבעלי אמונות.
תקציר מקורי באנגליתarXiv:2609.07943v1 Announce Type: cross Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like "belief" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpret
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית