יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למדנות רציונלית: כיצד ניתן לעודד אותה באמצעות ספורטיות קטנה

Extremely Sparse Supervision Incentivizes Reasoning Ability
למדנות רציונלית נעוצרת על ידי ספורטיות קטנה. חוקרים גילו שאפשר לעודד למדנות רציונלית באמצעות ספורטיות קטנה, כמו 0.05% מהטקסט. זה יותר דומה ללמידה טבעית.
תקציר מקורי באנגליתarXiv:2609.04565v1 Announce Type: new Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of all tokens. Surprisingly, this sparse supervision in most cases matches or surpasses full-token trai
קרא במקור המקורי