יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ללמוד: חידוש בבעיית הלמידה

Planning to Learn
חידוש חדש בבעיית הלמידה: הפסדי זיהוי נמוכים יותר עם 'הפסדי אורך-אפס'.
תקציר מקורי באנגליתarXiv:2610.03667v1 Announce Type: new Abstract: Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy
קרא במקור המקורי