כתבה
arXiv cs.AI ·
איזה תגים יש ללמוד? נקודת מבט של עריכה תגים לתפיסה מתמטית
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
מחקר חדש: עריכה של תגים לשיפור בהכשרה של תפיסה מתמטית
תקציר מקורי באנגליתarXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית