יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת תקינות בעת המבחן למודלי שפה גדולים

Test-time Calibration Learning for Large Language Model Reasoning
במאמר זה, המחברים מציגים פרקטיקה ללמידת תקינות בעת המבחן למודלי שפה גדולים. הם מציעים פרקטיקה חדשה שאינה תלויה בלייבלים, המאפשרת למודלים ללמוד תקינות בעת המבחן, כולל שמות מודלים/כלים/חברות.
תקציר מקורי באנגליתarXiv:2610.02695v1 Announce Type: new Abstract: Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time Cal
קרא במקור המקורי