כתבה
arXiv cs.AI ·
Diagnosing and Improving Probabilistic Reasoning in Large Language Models
תקציר מקורי באנגליתarXiv:2609.38005v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation f
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית