יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LLMs לומדות מטא-קוגניציה

LLMs learn different forms of metacognition when trained to predict their own accuracy
מודלי LLM מפתחים מטא-קוגניציה כאשר הם מאומנים לחזות את דיוקם. המחקר מראה שאימון כזה משפר את היכולת של המודלים להעריך את דיוקם.
תקציר מקורי באנגליתarXiv:2609.33886v2 Announce Type: replace Abstract: Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true
קרא במקור המקורי