יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למידת תקינות בזמן המבחן לתקשורת רציונלית של מודלי שפה

Test-time Calibration Learning for Large Language Model Reasoning
למידת תקינות בזמן המבחן לתקשורת רציונלית של מודלי שפה. המאמר עוסק בפיתוח של מודלי שפה שיכולים להציג ולהסביר את רמת הביטחון של התשובות שהם נותנים. המודלים הללו יכולים לשמש בתחומים שונים, כגון תקשורת רציונלית, חיפוש מידע ועוד.
תקציר מקורי באנגליתarXiv:2610.02695v1 Announce Type: cross Abstract: Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time C
קרא במקור המקורי