יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לימודי דימוי דקים מעודדים יכולת השכלה

Extremely Sparse Supervision Incentivizes Reasoning Ability
לימודי דימוי דקים מעודדים יכולת השכלה במודלי שפה. נמצא כי כ-0.05% מהטקסט הלמידה יכול להיות יעיל. ניתן להשוות זאת ללימודי דימוי מלא. נעשה שימוש במודל Qwen3 ובמודל GPT-5.
תקציר מקורי באנגליתarXiv:2609.04565v1 Announce Type: cross Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of all tokens. Surprisingly, this sparse supervision in most cases matches or surpasses full-token tr
קרא במקור המקורי