כתבה
arXiv cs.CL ·
Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models
תקציר מקורי באנגליתarXiv:2607.19608v1 Announce Type: new Abstract: Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts with their usual task behavior. We study this across three tasks - multiple-choice question answering (MCQA), sentiment classification, and mathematical question answering - by pairing a standard instruction with a conflicting non-standard one (select an incorrect option, output the opposite sentiment, or return twice the answer). This cross-task design allows us to test whether resistance to conflicting instructions is tied to specific task characteristics or reflects a broader behavioral tendency. As all predictions are scored against the original ground truth, a model that ignores the non-sta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית