כתבה
arXiv cs.AI ·
אבחון אבטלה לפני סקרים עם LLM
Before You Poll with LLMs: A Deliberative Diagnostic Framework
חוקרים פיתחו אבחון אבטלה ל-LLM, בודקים את GPT-5.1, Gemini 2.0 Flash, Claude Sonnet 4.5 ו-Llama 3.3 70B. האבחון מגלה כשלים בעדכון אמונות לאחר מידע חדש.
תקציר מקורי באנגליתarXiv:2609.15849v1 Announce Type: cross Abstract: Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית