כתבה
arXiv cs.CL ·
ללמוד לעצור בלי ללמוד לעצור
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
אימון עצמי-מפוקח של ביטחון משפר את יעילות ההיקש. מודלים כמו GPT ו-Qwen מאיצים היקש ב-25% ללא פגיעה בדיוק. השיטה משתמשת ב-600 בעיות אימון בלבד.
תקציר מקורי באנגליתarXiv:2609.31619v1 Announce Type: cross Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית