כתבה
arXiv cs.LG ·
Less Data Approximates More: Earning Faithful Confidence in High-Stakes Domains
תקציר מקורי באנגליתarXiv:2604.08454v2 Announce Type: replace Abstract: Large language models are increasingly deployed in high-stakes domains, where confident yet incorrect inferences may cause severe real-world harm, bringing the long-overlooked issue of confidence faithfulness to the forefront. A promising solution jointly optimizes unsupervised Reinforcement Learning from Internal Feedback (RLIF) with reasoning-trace-guided Reasoning Distillation (RD), yet it faces three persistent challenges, namely the scarcity of high-quality training corpora, factually unwarranted overconfidence, and erroneous updates amplified by indiscriminate fusion. Inspired by how human confidence accumulates from uncertainty to certainty, we propose Progressive Reasoning Gain (PRG) to measure whether reasoning steps progressivel
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית