כתבה
arXiv cs.CL ·
לפני שתצביעו עם LLMs: תפריט רפואי דלותי
Before You Poll with LLMs: A Deliberative Diagnostic Framework
חברת arXiv הכריזה על תפריט רפואי דלותי חדש שבודק את יכולת LLMs לעדכן את דעותיהם לאחר ידיעות חדשות. התפריט, המבוסס על תפריט רפואי דלותי, חושף כשלים שאינם נראים בבדיקות סטטיות. התפריט נבדק על חמש מודלי חידושים, כולל GPT-5.1, Gemini 2.0 Flash ו-Llama 3.3 70B.
תקציר מקורי באנגליתarXiv:2609.15849v1 Announce Type: new Abstract: Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית