יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ללמוד לעצור בלי ללמוד לעצור

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
אימון עצמי-מפוקח של ביטחון משפר את יעילות ההיקש. מודלים כמו GPT ו-Qwen מאיצים היקש ב-25% ללא פגיעה בדיוק. השיטה משתמשת ב-600 בעיות אימון בלבד.
תקציר מקורי באנגליתarXiv:2609.31619v1 Announce Type: cross Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoni
קרא במקור המקורי