כתבה
arXiv cs.CL ·
בדיקת יכולת LLM בהסקת היגום לוגי על מופעלי הסתברות
Benchmarking LLM Competence on Logical Inference over Probability Operators
חוקרים פיתחו בדיקת היכולת ל-LLM, בודקים 29 מודלים, רובם הראו העדפה לתשובות 'כן' או 'לא' ללא קשר לצורה הלוגית. רק 9 מודלים עברו את הסף האקראי.
תקציר מקורי באנגליתarXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negatio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית