כתבה
arXiv cs.AI ·
Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks
תקציר מקורי באנגליתarXiv:2610.02437v1 Announce Type: cross Abstract: Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styles. The tasks share an underlying semantic rule but differ in their prompt distributions and teachers' stylistic preferences. Using a tractable linear-softmax policy, we derive an exact decomposition of the updates into semantic and style components. We show that, at a common policy and prompt, SFT and RFT have parallel semantic updates but differ in their style dynamics. Starting from a policy with no within-class style preference, RFT with exac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית