כתבה
arXiv cs.AI ·
RL-ARC: איכון מודלים גדולים להיגיון
RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
RL-ARC הוא כלי לאיכון מודלים גדולים להיגיון. הוא משתמש בביטחון היגיון כאותות עזר לאיכון ביטחון תשובות. RL-ARC משפר את האיכון ומאפשר למודלים להעריך ביטחון באופן אדפטיבי.
תקציר מקורי באנגליתarXiv:2610.11352v1 Announce Type: new Abstract: Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-aware training methods for LMs, which incorporate objectives for uncertainty estimation into training, improve calibration but still exhibit overconfidence under distribution shift, while sacrificing reasoning performance. To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence. Specifically, RL-ARC leverages reasoning confidence as an auxiliary signal
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית