כתבה
arXiv cs.LG ·
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
תקציר מקורי באנגליתarXiv:2607.20526v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We present ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting. Confidence is elicited via prompting, requiring no access to model logits and making the framework applicable to both closed-source and open-source systems. The benchmark comprises 200 private multiple-choice questions across four categories: spatial reasoning, high-precision mathematics, word lookup, and unkno
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית