כתבה
arXiv cs.CL ·
קליברציה של מודלי שפה גדולים: מתגובה ליכולת
On Calibration of Large Language Models: From Response To Capability
במאמר זה, המחברים חוקרים את הקליברציה של מודלי שפה גדולים ומציעים פתרון חדש לבעיית ההערכה של תוצאות המודל. הם מציגים את פרקטיקת הקליברציה החדשה, המתמקדת בהערכת יכולת המודל לפתור שאלות, ומציעים דרכים לשיפור הקליברציה של המודלים.
תקציר מקורי באנגליתarXiv:2602.13540v2 Announce Type: replace Abstract: Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית