יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כאשר ביטחון לשוני וביטחון פנימי נפרדים במודלי שפה גדולים

When Linguistic and Internal Confidence Diverge in Large Language Models
מחקר: ביטחון לשוני של מודלי שפה גדולים נפרד מביטחון פנימי. המחקר חקר 30 מודלים ומצא שהביטחון הלשוני לא תמיד משקף את הביטחון הפנימי. המחקר גם מצא שמודלים שנלמדו על תווי הוראה נוטים לדווח על ביטחון גבוה יותר, אך גם על קפאון גדול יותר.
תקציר מקורי באנגליתarXiv:2608.28382v2 Announce Type: replace Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show hig
קרא במקור המקורי