יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

איזה תגים יש ללמוד? נקודת מבט של עריכה תגים לתפיסה מתמטית

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
מחקר חדש: עריכה של תגים לשיפור בהכשרה של תפיסה מתמטית
תקציר מקורי באנגליתarXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an in
קרא במקור המקורי