כתבה
arXiv cs.AI ·
When Linguistic and Internal Confidence Diverge in Large Language Models
תקציר מקורי באנגליתarXiv:2608.28382v2 Announce Type: replace-cross Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes sh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית